Any data scientists out there? What's your go to programming language and tools for your work?

rutrum@lm.paradisus.day · 3 days ago

Any data scientists out there? What's your go to programming language and tools for your work?

Kache@lemm.ee · edit-2 1 day ago

Hm, that’s kind of interesting

But my first reaction is that optimizations only at the “Python processing level” are going to be pretty limited since it’s not going to have metadata/statistics, and it’d depend heavily on the source data layout, e.g. CSV vs parquet

rutrum@lm.paradisus.day · 1 day ago

You are correct. For some data sources like parquet it includes some metadata that helps with this, but it’s not as robust at databases I dont think. And of course, cvs have no metadata (I guess a header row.)

The actually specification for how to efficiently store tabular data in memory that also permits quick execution of filtering, pivoting, i.e. all the transformations you need…is called apache arrow. It is the backend of polars and is also a non-default backend of pandas. The complexity of the format I’m unfamiliar with.