5 comments

  • vlovich1232 hours ago
    I feel like Debug should be lazily emitted altogether when it’s first used - just a special marker that’s never expanded since 99% of the Debug implementations aren’t used and having the rest marked #[cold] as inline is obviously wrong. Of course implementing it in practice sounds exceptionally difficult.<p>That being said, even the justifying performance improvement PR was itself a mix of improvements and regressions
  • Sharlin3 hours ago
    I’m fairly convinced that Debug should never be inlined. Display probably neither, the fmt machinery is heavy enough that not inlining is probably not a bottleneck even in serialization-heavy workloads. I’ve had to #[inline(never)] some of my own Debug&#x2F;Display impls, shrinking the binary by tens of kilobytes (out of a few hundred, so relatively a significant reduction).
    • aw16211072 hours ago
      For what it&#x27;s worth, according to the PR that added the annotation [0] doing so generally resulted in decreases in compile times and binary sizes on benchmarks. Furthermore, an additional experiment that avoided emitting the inline attribute on structs with &gt;5 fields resulted in benchmark regressions compared to always emitting the attribute [1]. I&#x27;d guess this is one of those things which may help in aggregate but hurts for specific cases.<p>That being said, one of the Rust devs indicated in the corresponding lobste.rs discussion [2] that they&#x27;re open to revisiting&#x2F;rebalancing things if they get enough bug reports indicating something is up, so it might not hurt to tag onto the bug report the author will (hopefully) eventually submit.<p>[0]: <a href="https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;pull&#x2F;117727" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;pull&#x2F;117727</a><p>[1]: <a href="https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;pull&#x2F;118031" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;rust-lang&#x2F;rust&#x2F;pull&#x2F;118031</a><p>[2]: <a href="https:&#x2F;&#x2F;lobste.rs&#x2F;s&#x2F;dldhpw&#x2F;rust_s_derive_often_implies_inline" rel="nofollow">https:&#x2F;&#x2F;lobste.rs&#x2F;s&#x2F;dldhpw&#x2F;rust_s_derive_often_implies_inlin...</a>
      • afdbcreid2 hours ago
        A reasonable conjecture was raised on lobsters that this is because `#[inline]` makes actual codegen (LLVM IR and down from MIR) lazy, and most `Debug` impls are never used.
        • infogulch56 minutes ago
          How much code is never used and compilation could be skipped entirely? Maybe applying a reachability pass to skip compiling unused code would be helpful.
          • mgsloan29 minutes ago
            A cross-crate dead code analysis would mean that compilation of a crate now depends on information about its dependents. This would break reuse of compiled crates and cause recompiles when the analysis changes.<p>Something does seem a little off about this, though. Ideally for this `Debug` case there would be an annotation that says &quot;compile this lazily, don&#x27;t inline&quot;. Maybe there doesn&#x27;t even need to be a new annotation, just `#[inline] #[cold]`. Which looks pretty weird, but might work already.
      • Sharlin2 hours ago
        Thanks, interesting!
  • scottlamb58 minutes ago
    I wonder if they ever considered a table-driven approach for `#[derive(Debug)]`, as `facet` [1] does. That would have been my first instinct for something this formulaic where binary size and compilation time matter more than execution speed. But my impression is facet hasn&#x27;t quite realized its promise on those fronts, so maybe the table-driven approach in std was similarly tried and rejected.<p>[1] <a href="https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;facet" rel="nofollow">https:&#x2F;&#x2F;crates.io&#x2F;crates&#x2F;facet</a>
  • zamazan4ik2 hours ago
    Or just try to avoid all of these optimization guesses by using Profile-Guided Optimization (PGO), that inserts&#x2F;deletes all inlines based on actual application runtime profile.
    • woodruffw2 hours ago
      The codebase in question (uv) uses PGO already. I suspect there isn’t a general way to guarantee that PGO ensures that only the “right” things get inlined.
  • api2 hours ago
    There&#x27;s a joke that the LLVM heuristic for whether to inline a function is &quot;return true;&quot; LLVM tends to inline aggressively.<p>You can control this behavior with opt level &quot;s&quot; or &quot;z&quot; or &quot;#[inline(never)]&quot;, but be aware that too little inlining can have large negative performance impacts.<p>It&#x27;s hard to get inlining exactly right without profile guided optimization.