3 comments

  • mbeavitt3 hours ago
    Why would someone want to use a nested function, practically speaking?
    • mananaysiempre3 hours ago
      Good C style is that every function that accepts a callback should also accept an opaque context pointer it then passes through unchanged to the callback. Usually the caller will allocate a structure on the stack or the heap, stash some of its local variables there, then use them in the callback. A nested function does the structure back-and-forth for you in the stack-allocated case. In GCC’s original formulation it also passes the context pointer implicitly<p><pre><code> size_t filter(bool (*predicate)(int), int *p, size_t n) { for (size_t r = 0, w = 0; r &lt; n; r++) { if (predicate(p[r])) p[w++] = p[r]; } return w; } size_t lowpass(int limit, int *p, size_t n) { bool lower(int value) { return value &lt; limit; &#x2F;&#x2F; use the parent&#x27;s local variable } return filter(lower, p, n); } </code></pre> but that requires an executable stack and TFA is about avoiding that part.
    • anta403 hours ago
      Say to strictly enforce modularity, e.g helper functions that can only be accessed within its function.<p>Pascal supports it (at least Turbo Pascal, no idea about ISO Pascal).
      • Joker_vD1 hour ago
        For a counterpoint, see David R. Hanson&#x27;s &quot;Is block structure necessary?&quot; (1981) [0] — back in those days, &quot;block structure&quot; meant nested routines with nested scopes — which argues that having instead a proper module system, with explicit control over what&#x27;s being exported from a module, not only gives a better modularity, decomposition, and encapsulation, but also simplifies both the language&#x27;s implementation, and the run-time structures it needs (remember displays, and the hardware support for them e.g. x86&#x27;s ENTER?).<p><pre><code> Block structure is traditionally considered an a priori requirement for algorithmic program- ming languages. Most new languages since Algol-60 have block structure. Reasons exist, however, to omit the general form of block structure — nested procedure definitions in which references to identifiers defined in outer procedures are permitted — from programming languages, especially those intended for systems programming applications. This paper reviews the concept of block structure and considers its advantages and disadvantages. It concludes that, in many cases, a module facility is superior to block structure and should be considered in lieu of block structure in future languages. </code></pre> [0] <a href="https:&#x2F;&#x2F;drh.github.io&#x2F;documents&#x2F;blockstructure.pdf" rel="nofollow">https:&#x2F;&#x2F;drh.github.io&#x2F;documents&#x2F;blockstructure.pdf</a>
      • kccqzy1 hour ago
        The traditional way of doing this in C is simply static functions. Every .c file has exactly one non-static function and all the other helper functions are static.
    • jcranmer3 hours ago
      When you want to use lambdas, but your language doesn&#x27;t have lambdas, so you reach for the nearest thing instead.
      • uecker2 hours ago
        Lambdas are just anonymous nested functions. But I like named nested functions more because they are more readable and would prefer them in most cases. Ideally you have both as most languages have.<p>I always wondered why C++ only added lambdas, but observing WG21 for a while, I assume this is just a random walk in language design. (not that it is different in WG14)
        • wasmperson41 minutes ago
          &gt; Lambdas are just anonymous nested functions.<p>The important feature of lambdas is that they are expressions, not that they lack a name. The advantage of function expressions is you can write the body of the function exactly at the place where it is used. With GCC nested functions you either have to write the body of the function before its first use or else write the declaration of the function twice.<p>This matters for long chains of continuation passing:<p><pre><code> foo(arg1, arg2, [](){ &#x2F;&#x2F; do some work bar(arg3, arg4, [](){ &#x2F;&#x2F; do some more work baz(arg5, arg6, [](){ }); }); }); </code></pre> Compare to the following, where the control flow is all out of order:<p><pre><code> void cb(void){ &#x2F;&#x2F; Do some work void cb2(void){ &#x2F;&#x2F; do some more work void cb3(void){ } baz(arg5, arg6, cb3); } bar(arg3, arg4, cb2); } foo(arg1, arg2, cb);</code></pre>
          • uecker34 minutes ago
            I agree with your point.<p>But I usually prefer the later anyway, because the code usually is not as nested anyway and having a name is often helpful, and also because I find the nested code with lambdas also not too readable. Other languages have better syntax for chaining functions in this way, i.e. with lambdas I would like to write like this:<p><pre><code> foo(arg1, arg2, _) .(int(int x)) { ... } .(int(int y)) { ... }; </code></pre> (edit: or something, I think I got it a bit wrong, but you get the idea)<p>But I agree, sometimes lambdas are better so it would be good to have both.<p>(There is the classical hack to define lambdas using statement expressions and nested functions.)
        • eru1 hour ago
          I can write numbers like three by just writing 3 in my code. When I want a named number I use a syntax like x = 3. Why should functions be any different? A language doesn&#x27;t need different ways to name things for each type of thing. Integers, strings, functions etc: they can all use the same mechanism for naming.
          • uecker55 minutes ago
            I agree if your language is designed like this from the beginning as functional languages are, but in C you already have different syntax for functions. (edit: rephrased)
      • mananaysiempre2 hours ago
        C++ bundles together a way to write functions inline in an expression (what I’d call “lambdas” in general) and a way to create closures with strictly nested lifetimes, but there’s no law of nature tying the two together. Even in C++ the essentially separate declaration “auto f = [&amp;](... blah ...) { ... 50 lines of code ... };” is pretty common. (And of course GCC’s nested functions predate C++11 by twenty years.)
    • psyclobe2 hours ago
      RAII style cleanup e.g. no gotos
    • sltkr2 hours ago
      For the non-capturing case: mainly to improve readability by allowing utility functions to be defined close to where they are used and with short names.<p>For the capturing case: to access context that is not available through global variables or function arguments, i.e., the same reason why closures are useful in other languages.<p>Here&#x27;s an example, where I have a list of points that I want to sort based on distance to a chosen target point. I can use qsort() which takes an arbitrary comparison function, but has no way to provide context to that function beyond the input arguments:<p><pre><code> #include &lt;stdio.h&gt; #include &lt;stdlib.h&gt; int main() { struct Point { int x, y; } points[3] = { { 3, 1 }, { 2, 2 }, { 5, 7 } }; struct Point target = { 4, 5 }; long dsq(const struct Point *p) { long dx = p-&gt;x - target.x, dy = p-&gt;y - target.y; return dx*dx + dy*dy; } int compare(const void *p, const void *q) { long a = dsq(p), b = dsq(q); return (a &gt; b) - (a &lt; b); } qsort(points, 3, sizeof(struct Point), compare); for (int i = 0; i &lt; 3; ++i) { printf(&quot;%d,%d\n&quot;, points[i].x, points[i].y); } } </code></pre> Note here that dsq() is a local function that accesses the `target` variable in the local function scope.<p>The usual workaround in standard C is to pass the necessary context as a function argument. That&#x27;s why qsort_r() exists, which takes a context argument to be passed to compare(), but that&#x27;s a non-standard GNU extension.<p>This practice of passing context pointers around is ubiquitous in C code, and it works, but it can get messy especially if you need access to multiple variables or variables from more than one nested scope. There is also a type safety issue: these context pointers are necessarily passed as void* which means they have to be cast back to the real type before use, which is where bugs can be introduced if the caller and receiver disagree on the actual type.
      • uecker44 minutes ago
        This is a good example. Without trampolines, this could look like this (Godbolt: <a href="https:&#x2F;&#x2F;godbolt.org&#x2F;z&#x2F;nK5fqMxjs" rel="nofollow">https:&#x2F;&#x2F;godbolt.org&#x2F;z&#x2F;nK5fqMxjs</a>).<p><pre><code> int main() { struct Point { int x, y; } points[3] = { { 3, 1 }, { 2, 2 }, { 5, 7 } }; struct Point target = { 4, 5 }; long dsq(const struct Point *p) { long dx = p-&gt;x - target.x, dy = p-&gt;y - target.y; return dx*dx + dy*dy; } typedef typeof(dsq) dsq_f; int compare(const void *p, const void *q, void *data) { wide(dsq_f) *dsq = data; long a = CALL(*dsq, (p)), b = CALL(*dsq, (q)); return (a &gt; b) - (a &lt; b); } qsort_r(points, 3, sizeof(struct Point), compare, &amp;CLOSURE(dsq_f, dsq)); for (int i = 0; i &lt; 3; ++i) printf(&quot;%d,%d\n&quot;, points[i].x, points[i].y); } </code></pre> There are slightly different ways how to define the helper macros, I am still experimenting a bit. Here you could avoid the typedef if defined differently. But ideally, there would be native language support that avoids these macros.
        • listeria15 minutes ago
          Well, if you&#x27;re already using qsort_r, what&#x27;s the point of using nested functions, if you can have a context pointer with the target?<p>And if you&#x27;re not using qsort_r, but reaching for _Thread_local, the target can be _Thread_local instead of dsq.
          • uecker2 minutes ago
            Fair. If you only have one object to access such as the target pointer, then it probably makes not much difference with void-pointer based APIs such as qsort_r (for new APIs it would add type safety). Where it removes more boilerplate code is when you have several such objects and would have to create an extra data structure to access them.
        • uecker40 minutes ago
          Or without qsort_r, you could use a thread local variable:<p><pre><code> typedef typeof(dsq) dsq_f; _Thread_local static wide(dsq_f) wdsq; wdsq = CLOSURE(dsq_f, dsq); </code></pre> <a href="https:&#x2F;&#x2F;godbolt.org&#x2F;z&#x2F;3e157c6b1" rel="nofollow">https:&#x2F;&#x2F;godbolt.org&#x2F;z&#x2F;3e157c6b1</a>
    • bobmcnamara3 hours ago
      Just a little cleaner than placing it in the global or file namespaces.
    • kloop3 hours ago
      So that you can name a section of code without polluting the namespace.
  • mananaysiempre4 hours ago
    What about your older patch where -fno-trampolines meant a function pointer could either be a code pointer or a closure (descriptor) pointer, distinguished by a tag?
    • uecker3 hours ago
      My old patch from 2018? This was not accepted to GCC because it relied on function pointers being aligned and there were concerns with this.<p>But I prefer this approach anyhow, as it does not impose any run-time cost for checking the tag, and is easier to optimize.
  • tpoacher3 hours ago
    What&#x27;s a &quot;trampoline&quot;?
    • jcranmer3 hours ago
      In this context:<p>Nested functions have a different ABI from regular C functions, due to the invisible static chain register that needs to be set up. C has no way of indicating this different ABI, so GCC happily lets you cast a nested function to a C function pointer by creating a little tiny function that puts the right value in the static chain register before calling the nested function. This little tiny function is the trampoline.<p>Since the trampoline needs to live somewhere, GCC puts it on the stack, requiring the stack to be executable and consequently a whole lot of people hate the feature because it&#x27;s a walking security nightmare.
      • uecker2 hours ago
        Correct (although the nightmare part is a bit exaggerated since return-oriented programming showed that non-executable stack does not help a lot). GCC can also put the trampoline on the heap, but this also has downsides.<p>For me the main downside of trampolines is that the optimizer can not de-virtualize the trampoline again. This could be implemented, but avoiding the creation of the trampoline in the first place is much better.
      • kccqzy1 hour ago
        C++ solves this problem by simply not allowing a nested function (lambda) to be converted to a function pointer, and thereby avoids this problem of trampolines and executable stack altogether. I think that’s a better design.
        • jcranmer49 minutes ago
          A C++ lambda that doesn&#x27;t close over anything <i>can</i> be converted to a function pointer: <a href="https:&#x2F;&#x2F;eel.is&#x2F;c++draft&#x2F;expr.prim.lambda#closure-12" rel="nofollow">https:&#x2F;&#x2F;eel.is&#x2F;c++draft&#x2F;expr.prim.lambda#closure-12</a> This feature does turn out to be useful if you need to pun a C++ interface into a function pointer for a C ABI function.
        • uecker52 minutes ago
          C++ has the same solution as I propose here: A wide function pointer type.<p>In C++ it is called std::function, but this comes with a bit of baggage. C++ 26 has std::function_ref which would be the exact equivalent to my wide pointer.<p><a href="https:&#x2F;&#x2F;godbolt.org&#x2F;z&#x2F;GaP9jb5rE" rel="nofollow">https:&#x2F;&#x2F;godbolt.org&#x2F;z&#x2F;GaP9jb5rE</a>
    • mananaysiempre3 hours ago
      Could be a number of things depending on context. In this case it’s a short function that adjusts some things and jumps to the actual functions (a “thunk” is another term for this). Specifically, if in GCC you write<p><pre><code> int f(int x) { int g(int y) { ... use x and y ... } ... h(&amp;g); ... } </code></pre> then what the compiled code for f does is construct <i>on the stack</i> a short piece of machine code:<p><pre><code> mov &lt;well-known register&gt;, &lt;frame pointer&gt; jmp &lt;start of g’s code&gt; </code></pre> and &amp;g points to the start not of g’s code but of this snippet on the stack, which has the parent function’s frame pointer compiled into it as a literal constant. The snippet is called a trampoline.
    • monster_truck3 hours ago
      It&#x27;s where you jump and then get immediately bounced back. Basically GOTOs with params
      • eru1 hour ago
        If you have proper tail call optimisation, then tail calls are GOTOs with params.<p>Trampolines allow you to simulate that, even when your compiler &#x2F; language doesn&#x27;t handle tail calls properly.
    • dgellow3 hours ago
      <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Trampoline_(computing)" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Trampoline_(computing)</a>