this post was submitted on 02 Aug 2024
344 points (97.5% liked)

Science Memes

11205 readers
2573 users here now

Welcome to c/science_memes @ Mander.xyz!

A place for majestic STEMLORD peacocking, as well as memes about the realities of working in a lab.



Rules

  1. Don't throw mud. Behave like an intellectual and remember the human.
  2. Keep it rooted (on topic).
  3. No spam.
  4. Infographics welcome, get schooled.

This is a science community. We use the Dawkins definition of meme.



Research Committee

Other Mander Communities

Science and Research

Biology and Life Sciences

Physical Sciences

Humanities and Social Sciences

Practical and Applied Sciences

Memes

Miscellaneous

founded 2 years ago
MODERATORS
 
you are viewing a single comment's thread
view the rest of the comments
[–] QuizzaciousOtter@lemm.ee 43 points 3 months ago (21 children)

Is 600 MB a lot for pandas? Of course, CSV isn't really optimal but I would've sworn pandas happily works with gigabytes of data.

[–] MoonHawk@lemmy.world 26 points 3 months ago* (last edited 3 months ago) (5 children)

What do you mean not optimal? This is quite literally the most popular format for any serious data handling and exchange. One byte per separator and newline is all you need. It is not compressed so allows you to stream as well. If you don't need tree structure it is massively better than others

[–] QuizzaciousOtter@lemm.ee 14 points 3 months ago

I think portability and easy parsing is the only advantage od CSV. It's definitely good enough (maybe even the best) for small datasets but if you have a lot of data you need a compressed binary format, something like parquet.

[–] elmicha@feddit.org 8 points 3 months ago

But which separator is it, and which line ending? ASCII, UTF-8, UTF-16 or something else? What about quoting separators and line endings? Yes, there is an RFC, but a million programs were made before the RFC and won't change their ways now.

Also you can gzip CSV and still stream them.

[–] merari42@lemmy.world 4 points 3 months ago (1 children)

Have you heard that there are great serialised file formats like .parquet from appache arrow, that can easily be used in typical data science packages like duckdb or polars. Perhaps it even works with pandas (although do not know it that well. I avoid pandas as much as possible as someone who comes from the R tidyverse and try to use polars more when I work in python, because it often feels more intuitive to work with for me.)

[–] driving_crooner@lemmy.eco.br 1 points 3 months ago

I used to export my pandas DataFrames as pickles, but decided to test parquet and it was great. It was like 10x smaller and allowed me to had the the databases on a server directory instead of having to copy everything to the local machine.

[–] candyman337@sh.itjust.works 1 points 3 months ago

If you have a csv bigger than like 500mb you need more than 8gb of ram to open it

[–] HappyFrog@lemmy.blahaj.zone 1 points 3 months ago

Wait till you hear about WSV

load more comments (15 replies)