Pandas

fossilesque@mander.xyz · 4 months ago

Pandas

Kausta@lemm.ee · 4 months ago

You havent seen anything until you need to put a 4.2gb gzipped csv into a pandas dataframe, which works without any issues I should note.

QuizzaciousOtter@lemm.ee · 4 months ago

I really don't think that's a lot either. Nowadays we routinely process terabytes of data.

Kausta@lemm.ee · 4 months ago

Yeah, it was just a simple example. Although using just pandas (without something like dask) for loading terabytes of data at once into a single dataframe may not be the best idea, even with enough memory.

QuizzaciousOtter@lemm.ee · 4 months ago

Is 600 MB a lot for pandas? Of course, CSV isn't really optimal but I would've sworn pandas happily works with gigabytes of data.

wallmenis@lemmy.one · 4 months ago

Just read a few at a time...

Barx [none/use name] · 4 months ago

And there are like 8 software projects dedicated to making pandas wrappers that work with large datasets because this is somehow better than engineers and statisticians learning SQL or some kind of distributed calculations strategy.

propter_hog [any, any] · 4 months ago

I do this daily haha

ColeSloth@discuss.tchncs.de · 4 months ago

"Constipated"