15 comments

  • finnborge3 hours ago
    This feels to me like a lovely example of both high-quality systems design and information theory. From my perspective, one of the more abstractly interesting thoughts surfaced by this write-up is the degree to which you&#x27;ve &quot;encoded&quot; a highly complex set of relationships through use of metaphor.<p>That type of encoding (starting with the statement: you&#x27;re at a teaching hospital) is profoundly more efficient than having to deeply and reliably articulate what each of the various roles your agents embody are, let alone their interactions.<p>Perhaps there is more room for any of us to consider what existing systems or identities are described abundantly in training sets and can be leveraged to ritualistically encode these social&#x2F;civic&#x2F;cultural dynamics saying: &quot;act as though you&#x27;re ___.&quot; Not exactly a new point, but one this write-up certainly is pulling me towards.<p>Thank you for sharing and the care you put into writing this!
    • baddash46 minutes ago
      i think role-play and utilizing the full power of language, stories, character, and narrative will unlock very sophisticated use-cases and in general a new dimension to agentic systems much in the way you&#x27;re describing.<p>probably what will work best in the future is specific training or fine-tuning against curated datasets of narrative fiction and&#x2F;or texts in general? not really sure.
    • ajstorm3 hours ago
      Thanks for the thoughtful response. We too were surprised by how deeply we could take the model of the teaching hospital to software engineering. Whenever we think to expand the system in one dimension or another, the teaching hospital model seems to have a nearby analogy.<p>I will confess however that some people internally find the model confusing. For example, one user couldn&#x27;t remember that to get an issue actioned, they needed to put it into the &quot;waiting room&quot;. They wanted to be able to action their issues without learning how complex medical systems operate. There seems to be more work on our plates to make this resonate with all users.
      • jacquesm2 hours ago
        Now imagine what it is like to run an actual hospital with real people whose lives (or the lives of their loved ones) are often at stake, with underpaid and overworked staff and with messy biological creatures as the subjects instead of bits and bytes. If anything this whole exercise should <i>also</i> give you a much deeper appreciation of the people that feel themselves called to help others.
      • jaggederest1 hour ago
        I had a similarly useful analogy of the legal system emerge, in a similar way. I think there&#x27;s a great analogy between policy in the legal system and policy in software engineering, and they have some awfully good (and very, very historical) ways to think about e.g. amending, repealing, and adjudicating things based on those policies.
    • slopinthebag1 hour ago
      yeah this is really cool. i&#x27;ve been thinking about sort of a mirrored idea, where you end up with wizards, mages, clerics etc and fantasy terminology used. idk if it would be as effective as this is, but maybe it would be more creative somehow?<p>i kinda want to try to build this off of github. it&#x27;s essentially just an event bus &#x2F; message queue that workers (agents) tap into.
  • mncharity53 minutes ago
    One role I didn&#x27;t see was patient advocate&#x2F;representative? That might be another approach to non-convergence - &quot;how is this going?&quot; and escalation.
    • ajstorm24 minutes ago
      We actually have a &#x2F;sinai-advocate skill, where a human can advocate on behalf of a stuck patient. We use it every once in a while when the labels get screwed up, or a workflow fails for some reason.
  • james_marks4 hours ago
    A hint at why GitHub actions have been unreliable. How many teams are running factories like this on GH infra now?
    • devin3 hours ago
      I posted in another threat about this: I am seeing a lot of people building their own little bespoke factories. They introduce endless quality gates until it slows development down, and then they add more agents to decide when to run certain actions, and on and on. The end result from what I&#x27;ve seen and personally participated in, is that it often winds up providing negative value in the software development lifecycle. It creates a whole lot of heat, but IMO is not helping the teams utilizing them to ship value any faster than they would with a more limited setup.
      • supermdguy1 hour ago
        I&#x27;ve gone through this cycle recently. I think static lint&#x2F;type checks are super useful, but agentic code review loops can easily go off the rails.
    • gchamonlive3 hours ago
      This is Microsoft scale, I&#x27;d be surprised if it made any difference these labs. It&#x27;s more likely it&#x27;s plain managerial mishandling of the infra in chasing new profit heights.
  • kinduff3 hours ago
    Since we are all building our version of this, what mistakes did you made until you arrived to a well-balanced solution? I really like that you know how much an issue is, because then you can start optimizing.
    • ajstorm2 hours ago
      I&#x27;d say that the first &quot;mistake&quot; we made was having it run in auto-merge mode. It was incredible to see what it could produce, and the speed with which it worked, but while the results seemed good, they were being produced at a rate that we couldn&#x27;t human-verify. This is not to say that they were bad, but that we had no way to convince ourselves that they were good.<p>Once we started to review the code in more detail (and put the hospital into &quot;human approval mode&quot; to stop the momentum) we found one major thing that we corrected.<p>When issues were decomposed into a DAG, sibling issues would often defer work that they thought that their sibling(s) would handle, and that work sometimes just got dropped. This was because there was no way for sibling issues to communicate with each other. We&#x27;ve since added a deferred scope ledger and policy. If any issue is going to defer scope it must comment about the deferral in the code, and write it to a special ledger with the rationale, and how&#x2F;when the deferral should be actioned. This has prevented several items from falling through the cracks, but we&#x27;re constantly refining what is a valid deferral and what should be fixed immediately.<p>There are several other things that we&#x27;ve corrected over time, which makes me think that another blog post is in order. The full list would be too extensive to try and address in the comments section.
      • aetherspawn1 hour ago
        Is a code comment and ledger the best way? Should the agent just fill out a form or something and attach it to the sub issue. This is how the hospital would work.
        • ajstorm14 minutes ago
          It often does that too, but we keep the ledger and the code comment as well, to ensure that if it doesn&#x27;t get resolved by the sibling, that it&#x27;s not lost. If the sibling does resolve it, it&#x27;s removed from the ledger and the code.
  • ajstorm1 day ago
    Rafi and I, who authored this post, will be hanging out here for any questions people may have.
    • losteric18 minutes ago
      are any of these artifacts public and available for inspection&#x2F;use?
    • Eridrus3 hours ago
      Did you do any ablation studies on what is actually useful vs what happens to just work because these systems can work around whatever people do?<p>I compare this to OpenAI&#x27;s symphony prompt which, at a high level, does exactly the same thing outlined here except model selection and doesn&#x27;t really have the need for roles or hospital metaphors.
      • ajstorm2 hours ago
        No, we haven&#x27;t performed any ablation studies yet - it&#x27;s a good suggestion.<p>The comparison with OpenAI&#x27;s Symphony is valid (the two systems do broadly the same thing). One thing that&#x27;s different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven&#x27;t validated that it does.<p>I confess that there&#x27;s a lot more validation we could do with this model but we just haven&#x27;t found the time. One thing we don&#x27;t cover in the post is that this is a side project for both of us, so we have less time than we&#x27;d like to perform experimentation and validation.
    • what30 minutes ago
      The blog post links to an issue in the Sinai repo, but it’s private or just doesn’t exist?
    • contingencies4 hours ago
      Which inherent limitations did you recognize in the metaphor before commencing this research?
      • ajstorm2 hours ago
        One thing that teaching hospitals have is a longer ladder of experience. For example many hospitals have medical students, interns, residents, chief residents, fellows and attendings. One idea we contemplated was to have &quot;treatment&quot; start on the cheapest possible model and see how it ended up, only escalating to more capable&#x2F;expensive models as necessary. We rejected this idea because we suspected that it would end up in more time&#x2F;tokens on the capable models to fix any problems.<p>In the end, medical students learn to become doctors. Cheaper models never learn to become more capable models, so the metaphor is not apt.
  • dingaling9113 hours ago
    Maybe I missed it, but I didn&#x27;t really see anything about the long term quality or maintainability of the code. All I see is agent agent agent.
    • ajstorm3 hours ago
      It&#x27;s true that we don&#x27;t have long-term maintainability data just yet. We&#x27;ve just shipped the first product of this model to customers and likely won&#x27;t have any detailed maintainability data for several months (and for good data, several years). We hope to author more blog posts on this experiment in the future.
      • Kinrany48 minutes ago
        Perhaps doing a random sample of the steps by hand will be a good way to notice maintainability issues?
  • jordanlewis1 day ago
    Great post!<p>One thing I&#x27;ve been experimenting with in my own agent orchestration system is an agent that hangs around and does post-merge acceptance testing after the work ships. Any plans to add a follow-up phase? Travel nurse?
    • ajstorm1 day ago
      That&#x27;s a very interesting idea. Sounds more like a &quot;routine follow-up&quot; in the medical model.
    • sglim1 day ago
      [dead]
  • zmj2 hours ago
    Nice writeup. Structured handoffs and external plan reviews are good takeaways.
    • ajstorm2 hours ago
      Thanks! And thanks for reading.
  • Veelox3 hours ago
    You give a very precise measure of redundancy in the skills. Can you give a bit more detail in how you decided you needed to audit them and how you went about it? Was it fully agent driven? Mostly human?
    • ajstorm3 hours ago
      All that credit goes to Rafi. I believe that he either noticed that they were getting long winded in a code review, or suspected that they needed trimming after reading this blog post from Anthropic: <a href="https:&#x2F;&#x2F;claude.dev&#x2F;blog&#x2F;the-new-rules-of-context-engineering-for-claude-5-generation-models&#x2F;" rel="nofollow">https:&#x2F;&#x2F;claude.dev&#x2F;blog&#x2F;the-new-rules-of-context-engineering...</a>.<p>It was definitely human driven, but I believe that the agents did the actual trimming.
      • Veelox2 hours ago
        That is helpful. Thank you :)
  • Spooky235 hours ago
    Reminds me of the “surgical team” development model in the Mythical Man Month.
    • sroerick2 hours ago
      I thought this too. It&#x27;s fun that they were working with DB2 as well.
  • git_rancher4 hours ago
    The patient “leaves” when the bug is fixed?
    • tough3 hours ago
      What would be the analogy if the patient dies?
  • ContinuityLab3 hours ago
    [flagged]
  • khotem4 hours ago
    [flagged]