Polars 2.0

(pola.rs)

381 points by simicd9 hours ago

23 comments

  • sureglymop3 hours ago
    I&#x27;ve used Polars before and can only recommend it.<p>It&#x27;s like you get a really good &#x27;query planner&#x27; like a DB would give you, but for your notebooks&#x2F;scripts&#x2F;etc. Much better than pandas imo.
    • jadbox2 hours ago
      Why use Polars over just postgresql or sqlite?
      • cle1 hour ago
        It is optimized for analytic workloads (default in-mem columnar layout), whereas PG and sqlite are OLTP DBs. It&#x27;s going to be insanely faster for those use cases.
        • zeristor27 minutes ago
          And with DuckDB?<p>DuckDB is an in process OLAP, I’ve been using it a lot, and I am keen to use Polars but DuckDB seemings to be flying for me.
      • Tuna-Fish1 hour ago
        Actual databases pay up to two orders of magnitude of speed for durability. If you can regenerate your dataset in case the DB drops it (or there is a power outage, or whatever), being able to complete a complex, write-heavy query in 10s instead of 15 minutes is actually very useful when doing analysis.
      • phillc731 hour ago
        If you want to use a database instead of Polars, for similar use cases, duckdb is a much better bet than postgressql or sqlite.
      • entropicdrifter2 hours ago
        Because those are databases? Polars is a data processing engine, not a database. They have different uses.
  • popularonion3 hours ago
    I worked on benchmarking in the past, including TPC benchmarks.<p>When you see a blog post like this, never interpret it as “database A is X% faster than database B”, there are just too many factors.<p>It’s more like “we put focused work into performance improvements and we expect certain workloads to perform better than the previous release”.<p>This seems like a good project, and benchmarking is a good way for a development team to iterate on performance. Just want to get my take out there.
    • andriy_koval3 hours ago
      But something like TPC is diverse enough to show whole picture?.. Especially compared to popular clickbench.
      • fphhotchips19 minutes ago
        It is absolutely without a doubt 100% not diverse enough.<p>I&#x27;ve spent years of my life running TPC benchmarks. They&#x27;re useful guidelines, but they&#x27;re simply not diverse enough to show the whole picture unless your picture is very simple. Certainly no single TPC benchmark in isolation. <i>Maybe</i> you can get a better idea by running all of them.<p>The largest gaps are around non-relational style data, JSON querying and the like. Nearly every organisation has <i>something</i> in that space now, and TPC-H&#x2F;TPC-DS don&#x27;t touch it at all.
        • andriy_koval13 minutes ago
          &gt; It is absolutely without a doubt 100% not diverse enough.<p>sure, nothing is 100%. But 90% may be good enough.<p>&gt; JSON querying and the like. Nearly every organisation has something in that space now, and TPC-H&#x2F;TPC-DS don&#x27;t touch it at all.<p>unnesting json into relational data is some trivial op, and then you come back into tpc-h&#x2F;ds realm.
      • orlp3 hours ago
        I think the person you&#x27;re replying to is more so saying that by choosing the right queries, machine, disk setup, caching, settings, thread count, RAM amount, etc, you can get quite different results. There is a reason everyone always wins their own benchmarks, and it doesn&#x27;t even have to be dishonest - you optimize and iterate for your own benchmark whereas everyone else just gets one shot to do well out of the box.<p>I tried my best to be as transparent and fair as possible, running everyone with out-of-the-box settings on a third party&#x27;s queries (DuckDB), a third party&#x27;s data generator (tpcgen-rs, from the DataFusion guys), on a stock setup available to everyone (AWS machines).<p>The one exception is that we also ran Polars locked to 32 threads on the large machine (in addition to the out-of-the-box setup), which was to highlight we <i>can</i> do a lot better on small data. We still suffer from a relatively high constant overhead on very high CPU count machines if the data isn&#x27;t large enough, but I&#x27;m working on fixing that. It&#x27;s possible that DuckDB &#x2F; DataFusion have similar scaling issues with high thread counts and would do better with 32 threads as well, I didn&#x27;t test that.
        • andriy_koval2 hours ago
          &gt; choosing the right queries, machine, disk setup, caching, settings, thread count, RAM amount, etc, you can get quite different results.<p>I think it is over-complication. TPC results used to be reported on some standard machines you can order, and now every one uses AWS metal for that. Its Ok if project has configs tuned to popular machine, I think it is representative approach.
  • gozzoo8 hours ago
    I&#x27;m not following the trends closely, but has Polars become a full replacement for Pandas? Are there use cases where one is better suited than the other?
    • SukadarBukadar3 minutes ago
      Pandas is better for slight in-place or per-row modifications, for loading from less conventional datasets, for transposition&#x2F;more nuanced row-based aggregation&#x2F;multi-axis manipulation, for performance when multiprocessing can be used, for interoperability with other libraries (e.g. plotting and statistics)... When you need to operate on huge datasets, use DuckDB, because its performance is even now on-par with Polars regarding speed, while handling huge or more complex joins is a huge win for DuckDB because it better offloads intermediate results to disk, while Polars just dies on me. <i>I haven&#x27;t tried such joins in Polars 2 though</i>
    • desipenguin7 hours ago
      From recent Python Bytes podcast (<a href="https:&#x2F;&#x2F;pythonbytes.fm&#x2F;episodes&#x2F;show&#x2F;496&#x2F;a-lake-house-in-seattle" rel="nofollow">https:&#x2F;&#x2F;pythonbytes.fm&#x2F;episodes&#x2F;show&#x2F;496&#x2F;a-lake-house-in-sea...</a>)<p>&gt; 1 Billion Row Challenge benchmark: Pandas took 4m28s vs. Polars 5.04s and DuckDB 5.19s — DuckDB also used 19x less memory<p>Python Vs Rust : In terms for speed - No comparison<p>(The above episode transcript has a link to blog post titled &quot;Pandas should go extinct&quot; )
      • jszymborski2 hours ago
        I&#x27;ve nearly entirely switched to DuckDB for anything more than like 500 or 1,000 rows or if there are a tonne of columns.<p>Polars is great, but I&#x27;m just too used to the Pandas API to use it as a replacement for the cases where DuckDB is overkill.
        • entropicdrifter1 hour ago
          That&#x27;s a shame, because Pandas has a really quirky&#x2F;legacy-burdened API and Polars is super clean by comparison. As someone who had Spark and Pandas experience before switching, Polars felt like Pyspark without the added mental overhead of needing you to think about multi-worker-node parallelism
          • jszymborski12 minutes ago
            It&#x27;s just muscle memor, I&#x27;ve been using Pandas daily for over a decade, but I&#x27;ll likely eventually switch over to Polars.<p>Part of the problem, as I said, is that I&#x27;m just spoiled by DuckDB when performance matters.
    • esco22927 hours ago
      Polars is effectively a full replacement for Pandas for 99.9% of all cases. The only exception I&#x27;m really aware of is if you&#x27;re working with geospatial data, as there isn&#x27;t yet a &quot;Geopolars&quot; equivalent of the commonly used &quot;Geopandas&quot;. However, Geopolars is still in active development and should eventually be production ready.
      • pattar2 hours ago
        There is a geospatial package for duckdb though which is pretty slick. It is actually how I first learned of duckdb. We were dealing with nationwide parcel datasets and need to apply transformations nationwide and save out to more parquet files. It was easier and cheaper to replace all of the pandas workloads with duckdb.
      • adeptima7 hours ago
        it’s on the correct path. i use rust for geo spatial and the gap with c, c++ closing rapidly or negligible in most cases<p>from <a href="https:&#x2F;&#x2F;github.com&#x2F;pola-rs&#x2F;geopolars&#x2F;tree&#x2F;main" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;pola-rs&#x2F;geopolars&#x2F;tree&#x2F;main</a><p>Comparison with GeoPandas<p>Imitation is the sincerest form of flattery! GeoPandas — and its underlying libraries of shapely and GEOS — is an incredible production-ready tool.<p>GeoPolars is nowhere near the functionality or stability of GeoPandas, but competition is good and, due to its pure-Rust core, GeoPolars will be much easier to use in WebAssembly.
        • ritchie466 hours ago
          Note that (for the time being), we started geopolars developement under: <a href="https:&#x2F;&#x2F;github.com&#x2F;pola-rs&#x2F;geopolars&#x2F;tree&#x2F;dev" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;pola-rs&#x2F;geopolars&#x2F;tree&#x2F;dev</a>
          • benrutter4 hours ago
            &gt; Note that (for the time being), we started geopolars developement under: <a href="https:&#x2F;&#x2F;github.com&#x2F;pola-rs&#x2F;geopolars&#x2F;tree&#x2F;dev" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;pola-rs&#x2F;geopolars&#x2F;tree&#x2F;dev</a><p>This is awesome!! I&#x27;d looked at the project only a month or so and it appeared abandoned, but I must have missed the off-main-branch development going on!
    • lmeyerov6 hours ago
      We were able to do a full port of GFQL from pandas to polars, cypher graph queries on dataframes, including both our CPU + GPU modes, and hit massive speedups: <a href="https:&#x2F;&#x2F;www.graphistry.com&#x2F;blog&#x2F;cypher-on-polars-cpu-gpu-graph-engine" rel="nofollow">https:&#x2F;&#x2F;www.graphistry.com&#x2F;blog&#x2F;cypher-on-polars-cpu-gpu-gra...</a><p>It&#x27;s been impressive!
    • niltecedu6 hours ago
      Yes and no, its not replacing the reason why pandas was popular ie data scientists, but it a full replacement of its pipeline usage, And I would saw also beating out spark
    • 3927 hours ago
      my understanding is Polars is faster, scales better without using external solutions, better API, +Rust. Pandas wins if you want to use what the vast majority of folks are using and have used in the past. Probably has a more complete set of helpers &#x2F; recipes for the little things you bump into when using it thoroughly, but in the age of LLMs, I think that&#x27;s minor.
      • vovavili7 hours ago
        &gt;Pandas wins if you want to use what the vast majority of folks are using<p>Vast majority of skilled developers are now using Polars, unless they are constrained by lack of Narwhals support in their third-party library of choice (e.g. Great Expectations, SHAP). That&#x27;s the more important trend to follow.
      • derriz6 hours ago
        It’s the API that gave me the push to leave Pandas. 10 or more years of occasional Pandas use and I still had to google for any non-trivial queries.<p>In that regard, I’m still waiting for a credible jq replacement…
        • bitbang4 hours ago
          Replacement for jq: <a href="https:&#x2F;&#x2F;github.com&#x2F;01mf02&#x2F;jaq" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;01mf02&#x2F;jaq</a>
        • boltzmann645 hours ago
          Learn SQL and interface with Duck. You will be 100x faster than Pandas&#x2F;Polaris duo at fraction of memory. Also SQL is supported literally everywhere with a much more capable than Pandas API. Duck outputs to a Pandas Dataframe, but just treat that like a dictionary. Do all your processing, filtering and aggregation in Duck.<p>Also, try fx.wtf as a replacement for jq. it comes with a in-built tui viewer that supports vi-keybindings. Ecmascript is built into fx.wtf so you can query the JSON with JS notation (where JSON was born). You can use any JS functions including map&#x2F;reduce&#x2F;filter or perform any kind of transformation instead of learning jq dsl that you will forget tomorrow.
          • entropicdrifter1 hour ago
            Polars also supports SQL now, per the announcement we&#x27;re all commenting on.
      • mhh__5 hours ago
        Avoiding pandas developers is a great reason to use polars imo
    • seemaze5 hours ago
      It has been for me. I greatly prefer the API, it fits my mental model much better. Give it a try!
    • minimaxir5 hours ago
      See &quot;Pandas should go extinct&quot;: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49668198">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49668198</a><p>tl;dr yes
      • entropicdrifter1 hour ago
        It kind of reminds me of Dungeons and Dragons, in a sense. Pandas is so popular that people learn&#x2F;use it because of its popularity, rather than because it&#x27;s the best for any one use-case. Polars is cleaner, faster, and can easily be scaled up for production use-cases so your little POC script for local analysis can be productionized really easily, but there are holdouts still on Pandas because it was so complicated to learn with so many little extra rules to learn to avoid paper cuts that they feel like learning another data manipulation tool would be really hard.<p>DnD does the same thing: it&#x27;s the most popular but its rules are this awkward hybrid of legacy cruft and some modern ideas, so learning it a huge effort, which means most people who play it aren&#x27;t willing to try any other RPG systems even though most of them are <i>dramatically</i> easier to learn because they were built with a clean design from the ground-up.<p>It&#x27;s the sunk-cost fallacy as applied to learning something complex, combined with something like the horn effect (inverse of the halo effect) making any competitors look equally complex even if they&#x27;re not, causing long-time Pandas users&#x2F;DnD players to strongly resist even looking at other options. Basically the frustration of learning these older systems seemingly traumatizes some people into never straying.
    • bmitc6 hours ago
      There are awkward things. For example, if you ingest a nanosecond resolution timestamp, there&#x27;s no way to re-export that out of the Polars dataframe with nanosecond resolution.
      • simon-b44 minutes ago
        I&#x27;m not sure this is true e.g. you can specify schema `pl.Datetime(&quot;ns&quot;)`. That survives roundtrip to&#x2F;from parquet in my experience. True though that the default is `us` whereas pandas defaults to `ns`.<p><a href="https:&#x2F;&#x2F;docs.pola.rs&#x2F;api&#x2F;python&#x2F;stable&#x2F;reference&#x2F;expressions&#x2F;api&#x2F;polars.datetime.html" rel="nofollow">https:&#x2F;&#x2F;docs.pola.rs&#x2F;api&#x2F;python&#x2F;stable&#x2F;reference&#x2F;expressions...</a>
      • knottn59 minutes ago
        Feels like a feature to me. You could just keep the original timestamp as a string.
    • gremlinunderway1 hour ago
      From experience I&#x27;ll say one thing that Polars doesn&#x27;t have great interfacing with is doing things like string concatenation, like taking multiple columns and combining them in with static string text in complex ways to create new columns.<p>Pandas has a really simple ability to just define a new column with<p>`df[&#x27;col_a&#x27;] + &quot;text&quot; + df[col_b&#x27;]` where &quot;text&quot; can be any string text inbetween your column values from col_a and col_b<p>If i remember correctly while you can do pl.col(&quot;col_a&quot;) + pl.(&quot;col_b&quot;) for plain concatenation, you can&#x27;t mix in static text strings like you can with pandas and I haven&#x27;t found really elegant ways to do that personally. Whereas I&#x27;ve found polars doesn&#x27;t have as simple of a way to do that. You can choose to add one separator and make that anything you want, but only one separator and only inbetween the two values (so no suffixes or prefixes for example).<p>That being said, I hate everything to do with the pandas API (especially with its indexing system) and really prefer the more polars API for anyone coming from a SQL or database background. Pandas really shows its sort of academia background rather than a data engineering origin.
      • orlp1 hour ago
        I&#x27;m not sure what you tried. Is this not what you want?<p><pre><code> &gt;&gt;&gt; df = pl.DataFrame({&quot;x&quot;: [&quot;a&quot;, &quot;b&quot;, &quot;c&quot;], &quot;y&quot;: [&quot;d&quot;, &quot;e&quot;, &quot;f&quot;]}) &gt;&gt;&gt; df.with_columns(new=pl.col.x + &quot; text &quot; + pl.col.y) shape: (3, 3) ┌─────┬─────┬──────────┐ │ x ┆ y ┆ new │ │ --- ┆ --- ┆ --- │ │ str ┆ str ┆ str │ ╞═════╪═════╪══════════╡ │ a ┆ d ┆ a text d │ │ b ┆ e ┆ b text e │ │ c ┆ f ┆ c text f │ └─────┴─────┴──────────┘</code></pre>
  • tomrod4 hours ago
    Well done, Polars team!<p>Everything that I build greenfield moving forward I plan to use DuckDB, Polars, or PyArrow. Pandas was a great grandfather of a project (I actually cut my OSS contrib teeth on it, how the time flies)! I&#x27;ll always appreciate the improvement pandas brought over SAS.
    • genxy4 hours ago
      If Pandas was the great grandfather, R data.frame is the great-great grandfather. R data.frames directly inspired Pandas.
      • tomrod2 hours ago
        Agreed. For its time, R was a lot of fun to play with data pipelines, visuals, Quarto, and frontier stats. It&#x27;s a shame it&#x27;s so hard to make it work for a large swath of production use cases.
    • sanderjd3 hours ago
      Do you have a take on when each of these three choices is the best one? I totally agree that these are the good choices, but I still find myself hesitating about which thing to reach for when!
      • tomrod2 hours ago
        Depends on need. We started using PyArrow on a reporting microservice when we realized we needed no additional functionality that pandas provided since it has better data type ergonomics. DuckDb is a great go-to for SQL based transformations when working with parquet files outside a managed system like Databricks. I want to actually test duckdb versus polars with a few lower level places like iceberg on S3.
        • sanderjd2 hours ago
          Yeah makes sense. The SQL thing has also been my differentiator (and yeah, straight up arrow if you aren&#x27;t doing much transformation of the data), but now I&#x27;m curious whether polars sql might be just as good. I kind of like that duckdb allows me to work with a database file, like a sqlite db. But maybe persisting parquet (or arrow directly?) is just as good?<p>This is why I asked someone else who is also figuring this out!
          • tomrod1 hour ago
            Really comes down to query speed and management cognitive cost. I look forward to trying the new version for polars.
  • therno7 hours ago
    I use Polars 2.0(rc) to (pre)calculate billions of weather scores on <a href="https:&#x2F;&#x2F;therno.com" rel="nofollow">https:&#x2F;&#x2F;therno.com</a> and it has been a lifesaver<p>Happy that I can upgrade to 2.0 final tonight.
  • niltecedu6 hours ago
    A bit surprised about the datafusion results from the post, I have tried it time and time again, but datafusion has always been the leading&#x2F;trading blowers with polars for our workfloads with duckdb being vastly slower.
    • f311a6 hours ago
      There are some benchmarks for the previous version <a href="https:&#x2F;&#x2F;benchmark.clickhouse.com&#x2F;#system=+ti%20rud|Dusa,s|PDa|olrP&amp;type=-&amp;machine=-6t|ca2|6ax|g4e|6ale|3al&amp;cluster_size=-&amp;opensource=-&amp;hardware=+c&amp;tuned=+n&amp;storage=-&amp;metric=combined&amp;queries=-" rel="nofollow">https:&#x2F;&#x2F;benchmark.clickhouse.com&#x2F;#system=+ti%20rud|Dusa,s|PD...</a>
      • orlp6 hours ago
        Doing the benchmarks for 2.0 on the large AWS metal machines at small data sizes (SF=10) really opened my eyes that we have some low-hanging fruit in Polars when it comes to optimizing our constant overhead for smaller queries.<p>For example our join currently does a full partition into T partitions, for each of the T threads. Overall we create T^2 partitions, which on a 192-core machine is non-trivial. Great if you have a ton of data to feed that with, but if you &#x27;only&#x27; have a few dozen million rows it becomes rather small. This is the primary reason we saw in the benchmarks that Polars pinned to 32 threads beats 192 thread Polars at SF=10.<p>I&#x27;ll be working on improving that soon. I expect that to have a big impact on SF=10, and a decent impact on ClickBench, which sits between SF=10 and SF=100 in terms of rows.
  • sgarland8 hours ago
    TIL that Polars supports SQL. Amazing.
    • brap5 hours ago
      At what point can we say that Polars is basically an in-memory database? (Genuine question)
      • sanderjd3 hours ago
        I think it always has been? All of these data frame libraries are very much akin to olap databases.
    • dist-epoch5 hours ago
      It did since the first releases, but it was limited.<p>Now it seems they want to go head to head with DuckDB.
      • efromvt5 hours ago
        Love competition in the local data SQL space, makes everyone better
  • PLenz1 hour ago
    Does polars have a good equivalent of geopandas yet?
  • raoulj1 hour ago
    Looks like SQL does not support `asof_join()`, not sure what else it doesn&#x27;t support.
    • orlp1 hour ago
      Polars does have an asof join so I guess the mapping hasn&#x27;t been made yet from SQL. Could you open up an issue at <a href="https:&#x2F;&#x2F;github.com&#x2F;pola-rs&#x2F;polars&#x2F;issues" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;pola-rs&#x2F;polars&#x2F;issues</a> ?
      • raoulj58 minutes ago
        I use asof join frequently in polars! Happy to open the issue.
  • jdefting3 hours ago
    I would be interested in seeing memory usage differences in these benchmarks. I’ve had issues with excessive memory usage in polars compared to DuckDB.<p>I’m guessing most of the this disparity should be solved by the steaming engine.
    • kaathewise3 hours ago
      Polars relies on threading heavily, even when streaming files. And it appears that each thread loads quite a bit of memory. I&#x27;ve encountered OOM issues when incrementally reading Arrow IPC files which had very large batches. Fixed it by setting $POLARS_MAX_THREADS to 1, which amusingly also improved the performance on my very narrow task.
    • orlp3 hours ago
      Peak memory usage per query is in the raw data (in `results&#x2F;`) in the repository: <a href="https:&#x2F;&#x2F;github.com&#x2F;pola-rs&#x2F;polars-2.0-benchmark&#x2F;" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;pola-rs&#x2F;polars-2.0-benchmark&#x2F;</a>.
  • dkgs8 hours ago
    A coincidence with the fact duckdb is supposed to release 2.0.0 very soon? :)
    • orlp8 hours ago
      Actually yes. We had already been planning to do 2.0 for a long time. We originally said we&#x27;d move on from 1.x quickly when released 1.0 but ended up staying at 1.x much longer than intended.<p>From a quick check our first PRs were merged to the 2.0 branch in June:<p><pre><code> 2026-06-17T21:27:51Z #27993 chore: Stop coercing `pl.col(...)` to selector ... 2026-06-18T14:19:26Z #27996 chore!: Replace multi-seed hash API with a single seed 2026-06-19T07:05:11Z #27991 chore(python!): Remove `Expr.flatten` function</code></pre>
    • m00dy8 hours ago
      Are they competing for something?
      • vindex107 hours ago
        &gt; first class SQL support, which together with the performance improvements has Polars leading DataFusion and DuckDB in TPC-H and TPC-DS1 benchmarks,<p>Apparently about something ))
  • hnd9q09qk48 hours ago
    Thing I care about most is whether the old eager-vs-lazy footguns got cleaned up. Half my bugs were a stray collect() in a loop killing the query plan.
  • dartharva7 hours ago
    Been using polars for over a year now, it is fantastic.
  • breezybottom3 hours ago
    Looks like the Claude skill hasn&#x27;t been updated
  • Kinrany7 hours ago
    How does Polars relate to DataFusion these days? There&#x27;s no reason for them not to converge into a single ecosystem, is there?
    • esafak7 hours ago
      It does not use datafusion. <a href="https:&#x2F;&#x2F;github.com&#x2F;pola-rs&#x2F;polars&#x2F;issues&#x2F;6197" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;pola-rs&#x2F;polars&#x2F;issues&#x2F;6197</a>
      • Kinrany2 hours ago
        Yes, but why?<p>I mean, it&#x27;s very clear that the projects have a lot in common, even if building one on top of another directly is not a good idea.
        • esafak3 minutes ago
          In the issue, polars&#x27; creator says it is to reduce dependencies:<p><pre><code> I don&#x27;t want polars to depend on on the datafusion crate, that is a complexity that is not needed.</code></pre>
  • patrick_rtk2 hours ago
    amazing project !
  • Centigonal5 hours ago
    out of core sounds awesome! biggest thing that forced me to switch from pandas&#x2F;polars to other solutions back in the day.
  • bjourne3 hours ago
    Pandas is amazing, Polaris is amazinger!
  • Vaslo7 hours ago
    My team is all moving over to polars and DuckDB
    • matthewpick6 hours ago
      What was your previous setup &#x2F; stack?
  • yshvrdhn4 hours ago
    what about daft project ?
  • jt-s5 hours ago
    Although I find pandas a bit aggravating in many ways, for myself and my equally idiotic laboratory scientist pals, seems that it is the default way you might interface with other libraries like SciPy (i.e. they expect things as NumPy arrays or pandas dataframes). Is this a real issue or will most things happily accept a polars dataframe? We don’t work with such large datasets that speed is likely a huge concern tbh.
    • 0cf8612b2e1e3 hours ago
      One nice development in this space is the narwhals library - it is a dataframe agnostic library. It allows you to seamlessly switch between pandas, polars, modlin, or any of the variations coming out.<p>Narwhal is still fairly new, but I expect its usage to spread since most packages only require rudimentary dataframe manipulation (set a value, math been these two columns, etc) where the limited api surface is not a problem.<p>Narwhals is also a much cheaper dependency to add than polars&#x2F;pandas&#x2F;etc so it is a somewhat easy sell to incorporate.
    • niksmather4 hours ago
      You can convert to numpy using .to_numpy().<p>It&#x27;s also got much better support for more complex array shapes (e.g. each row storing an array). At least it did last time I used pandas!
  • tonyhart78 hours ago
    finally long time coming<p>cant wait to upgrade my Quant trading bot
    • geodel3 hours ago
      Not to take away any prop knowledge from you but I&#x27;ve been playing with some very very early ideas of getting some market data, save in duckdb do some analysis, create some kind of portfolio and buy&#x2F;sell via API and nowhere close implementation. Yours seems rather established platform. Would you be able to share kind of high level architecture of bot?
      • tonyhart72 hours ago
        use a higher timeframe to understand the market phase, and a lower timeframe for entries.<p>when I started, I tried mimicking a human trader. if you want to expand into full quant, be aware of overengineering (past mistake of mine).<p>you can copy or mimic institutional desk strategies. most of the concepts are fine, but the devil is in the details.<p>sometimes you open too early or too late. finding that sweet spot, where you don’t want a lot of drawdown, requires a lot of fine-tuning.
    • m00dy7 hours ago
      I&#x27;m also using it for BlockRotate, it&#x27;s time to upgrade.
      • tonyhart74 hours ago
        I mainly trade xauusd and fx<p>more predictable
  • mrtimo6 hours ago
    I can use pandas to clean a dataset, but each cleaning task is usually one line of code. OTOH, With DuckDB with one SQL statement I can replace 40+ lines of polars&#x2F;pandas. You may reply, SQL isn&#x27;t as easy to understand! Fair point, it&#x27;s a declarative language... which is why I use Malloy. Malloy is to TypeScript as Javascript is to SQL. Malloy is much easier to read and write (just as TypeScript is) because it has a built in semantic model -- all the joins, measures, and dimensions are done in one place.<p>Here is an example [1] of visualizing college football games. Here are all the queries, and semantic model that power all the visualizations [2] Here is the AI generated typescript&#x2F;react that does the visualizations [3]. The Malloy ecosystem has Malloyyo and Publisher which are replacements for PowerBI and Tableau and Looker. Here is another example for visualizing global trade [4].<p>[1] - <a href="https:&#x2F;&#x2F;mrtimo.github.io&#x2F;cfb-games&#x2F;games-2026.html?week=Week+5&amp;sort=thrill" rel="nofollow">https:&#x2F;&#x2F;mrtimo.github.io&#x2F;cfb-games&#x2F;games-2026.html?week=Week...</a> [2] - <a href="https:&#x2F;&#x2F;github.com&#x2F;mrtimo&#x2F;cfb-games&#x2F;blob&#x2F;main&#x2F;drives.malloy" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;mrtimo&#x2F;cfb-games&#x2F;blob&#x2F;main&#x2F;drives.malloy</a> [3] - <a href="https:&#x2F;&#x2F;github.com&#x2F;mrtimo&#x2F;cfb-games&#x2F;blob&#x2F;main&#x2F;dashboards&#x2F;games-2026.tsx" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;mrtimo&#x2F;cfb-games&#x2F;blob&#x2F;main&#x2F;dashboards&#x2F;gam...</a> [4] - <a href="https:&#x2F;&#x2F;tradeexplorer.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;tradeexplorer.org&#x2F;</a>
    • entropicdrifter1 hour ago
      The linked announcement includes the facts that Polars now supports SQL and the new version scales better than DuckDB