54 comments

  • HarHarVeryFunny9 hours ago
    The summary &quot;There are still clear limits. Gemini 3.5 Flash remains a better practical choice [than GPT 5.6 Sol] for high-volume detection and counting in our benchmark, especially at its price.&quot; seems rather understated !<p>GPT 5.6 Sol was outperformed on all benchmarks by Gemini 3.5 Flash, apart from a single exception (OCR) where Fable was the winner.<p>Gemini 3.5 Flash not only outperformed GPT 5.6 Sol, but did so at 1&#x2F;3 of the cost.
    • SkalskiP8 hours ago
      Hi, I’m the author of this blog post. I wrote it about 4 weeks ago, and the VLM world is moving so fast that it’s already kinda outdated. I think Gemini 3.7 Flash might be a better choice now, especially when you factor in the price.<p>Here’s a comparison of the best low-cost models I put together last week. What’s crazy is that Gemini 3.7 Flash is now 50% off on OpenRouter, and this chart doesn’t even account for that discount. <a href="https:&#x2F;&#x2F;x.com&#x2F;skalskip92&#x2F;status&#x2F;2088032652301304121?s=20" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;skalskip92&#x2F;status&#x2F;2088032652301304121?s=20</a>
      • MostlyStable7 hours ago
        Curious why you didn&#x27;t try Gemini 3 pro? That is the model I&#x27;ve been using for OCR entry of handwritten datasheets (JPGS of datasheets, structured JSON output). At my scale, the cost of 3 pro is basically not an issue, but if there are improvements in quality, I&#x27;d definitely be willing to explore other models
        • gdudeman4 hours ago
          In my experience starting with Gemini 2.5 Pro, moving to 3 and 3.1, 3.5 Flash, 3.6 Flash, and finally 3.7 Flash, 3.7 Flash is just as good if not better than 3 especially on high resolution mode (same token count per page as 3.1).<p>I run complicated, messy PDFs through these models. 2.5 Pro required a lot of kludgy hacks to get it to fully &quot;see,&quot; but from 3.1 pro on I&#x27;ve removed many of them and haven&#x27;t spotted problems.<p>3.7 Flash scores better than 3.1 pro on most benchmarks, leading me to believe that even if your OCR requires reasoning to interpret text or data, 3.7 Flash is probably going to be better.
        • bastawhiz5 hours ago
          3 Pro is quickly approaching one year old. There&#x27;s almost no reason to benchmark it, especially since a new version of Gemini Pro was supposed to be released mid 2026 and hasn&#x27;t seen the light of day.
          • MostlyStable5 hours ago
            That would make sense if we already knew that, for these kinds of tasks it was significantly worse. The tests that I&#x27;m aware of for these tasks show it as still performing near the top.
          • tziki3 hours ago
            I think it definitely makes sense since it&#x27;s still the best Google has to offer in the &quot;pro&quot; tier.
            • bastawhiz3 hours ago
              3 and 3.1 Pro are both marked as deprecated by Google. Even if they&#x27;re the best Google offers, it would be foolish to choose a model that&#x27;s explicitly deprecated.<p>It&#x27;s not a technical problem, it&#x27;s a commercial one. If Google can&#x27;t ship a model to replace the one they deprecated, that tells you everything you need to know about choosing a Gemini model for whatever you&#x27;re trying to do.
              • heaney-5552 hours ago
                3.1 Pro is not deprecated!
                • bastawhiz47 minutes ago
                  <a href="https:&#x2F;&#x2F;ai.google.dev&#x2F;gemini-api&#x2F;docs&#x2F;deprecations" rel="nofollow">https:&#x2F;&#x2F;ai.google.dev&#x2F;gemini-api&#x2F;docs&#x2F;deprecations</a><p>That link shows 3.1 pro listed as deprecated with no replacement model.
        • yieldcrv4 hours ago
          The “pro” moniker means nothing<p>these models aren’t successors and barely have a common ancestor, they are independently baked in the training oven and assigned a semantic version randomly by someone trying to show initiative but not trying to do on the toes of the last guy who got promoted first<p>So 3 pro is outdated and will likely never exit preview<p>The “flash” and “lite” models are the real “pro” in colloquial ideas of fleshed out and capability, at this point.<p>they’re better, faster and cheaper, larger context windows keeping up with the industry and more
          • heaney-5553 hours ago
            They are smaller models, and you can tell. Small models make dumb common-sense mistakes that big models never do. This is the &quot;smell&quot; many talk about.
            • sidibe3 hours ago
              Do you have cases where you still see 3.1 pro outperforming 3.7 flash?
              • heaney-5552 hours ago
                Yes, for complex questions of biology, physics, and analysis of anomalies.<p>3.7 Flash is better at coding, sure, but AI is not just for coding.
            • yieldcrv3 hours ago
              hasn&#x27;t been an issue since 3.5 for me, what have you seen, say, in the last two months
              • heaney-5552 hours ago
                For complex questions of biology, physics, and analysis of anomalies, 3.1 Pro is still better than 3.7 Flash for me.<p>3.7 Flash is better at coding, sure, but AI is not just for coding.
      • Melatonic4 hours ago
        What about Gemma ?
    • ImageXav5 hours ago
      Gemini tops their vision evals [0] by a mile, with 4&#x2F;5 top spots going to variants of it. Qwen is the only other contender, likely due to how good it is for object detection, where it crushes the competition [1].<p>[0] <a href="https:&#x2F;&#x2F;playground.roboflow.com&#x2F;evals" rel="nofollow">https:&#x2F;&#x2F;playground.roboflow.com&#x2F;evals</a><p>[1] <a href="https:&#x2F;&#x2F;playground.roboflow.com&#x2F;evals&#x2F;object-detection" rel="nofollow">https:&#x2F;&#x2F;playground.roboflow.com&#x2F;evals&#x2F;object-detection</a>
    • MrBuddyCasino9 hours ago
      Yeah I was thinking about giving Luna a go with my PDF data extraction, but I think I‘ll stay on Gemini. It does a very good job.
      • bicx9 hours ago
        Gemini is still my top choice within production software for typical data extraction from unstructured data. Gemini Flash Lite feels like a cheat code for speed, and it&#x27;s really cheap.<p>Some other Chinese models are also fast and cheap, but a harder sell in a U.S. production environment.
        • ComputerGuru5 hours ago
          Speaking from experience here, flash lite models have amazing price, speed, and perform far above their size, but are susceptible to very bad instruction following and recall when either complexity or context size inch up. They’ll just forget to apply your instructions to portions of the input, and repeat parts of the input that should be returned verbatim as direct quotes but with subtle changes (breaking urls, for example).
          • MrBuddyCasino4 hours ago
            Yes you have to continuously tune the prompts ever so subtly. 3.5 is a lot better than than 3.1 tho.<p>Important to remember that json schema instructions take precedence over the normal prompt, so move as much into property descriptions as possible.
            • ComputerGuru4 hours ago
              This was 3.5 flash lite, actually, and after prompt tuning. It was very clearly an issue that correlated with input (JSON array) size, the more elements in the batch, the higher the error rate.<p>3.0 flash (not lite) handled it like a champ though, fwiw.
        • MrBuddyCasino8 hours ago
          Yeah Gemini 3.5 Flash Lite is really good. Which Chinese models can you recommend?
          • SkalskiP8 hours ago
            Hi, I’m the author of this blog. It depends on how strong of a model you need, but in general, Qwen is easily the best among the Chinese models right now.<p>Over the last two weeks, Qwen released two new models. Qwen3.8-Max is totally insane, but it’s only available through the Alibaba Cloud API. I wrote a similar blog covering Qwen3.8-Max: [<a href="https:&#x2F;&#x2F;blog.roboflow.com&#x2F;qwen3-8-max&#x2F;" rel="nofollow">https:&#x2F;&#x2F;blog.roboflow.com&#x2F;qwen3-8-max&#x2F;</a>](<a href="https:&#x2F;&#x2F;blog.roboflow.com&#x2F;qwen3-8-max&#x2F;" rel="nofollow">https:&#x2F;&#x2F;blog.roboflow.com&#x2F;qwen3-8-max&#x2F;</a>)<p>If you’re looking for something you can run locally, Qwen3.8-27B might be a great option. On Friday, I did a quick comparison between Qwen3.8-Max and Qwen3.8-27B: [<a href="https:&#x2F;&#x2F;x.com&#x2F;skalskip92&#x2F;status&#x2F;2088411215441621469?s=20" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;skalskip92&#x2F;status&#x2F;2088411215441621469?s=20</a>](<a href="https:&#x2F;&#x2F;x.com&#x2F;skalskip92&#x2F;status&#x2F;2088411215441621469?s=20" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;skalskip92&#x2F;status&#x2F;2088411215441621469?s=20</a>)
            • kanemcgrath6 hours ago
              Googles local gemma models which target roughly the same parameter count range, are known for being a lot better at vision tasks than qwen, no idea if 3.8 has changed that though
              • SkalskiP5 hours ago
                Really? Gemma4-31B should be better than Qwen3.8-27B? I&#x27;m happy to test that.
          • b3458 hours ago
            I&#x27;ve been using Qwen3.5-9B, hosted locally for PDF data extraction and it performs pretty well when extracting data from tables and infographics
        • msp268 hours ago
          [dead]
      • dannyw7 hours ago
        Gemini is honestly an excellent LLM with many capability strengths.<p>For example, 3.7 Flash is #1 on MMLU Pro and AA’s agentic spreadsheets&#x2F;docs benchmark, etc. Yes, beating Fable.<p>Agentic coding is only one dimension.
        • fau6 hours ago
          Anecdotally, Gemini Flash is the leader for a particular use case of mine and has been since at least version 2.5. But now there&#x27;s also Luna as the first real competitor thanks to the price cut.<p>My worry is that this is a zero-sum game and when Gemini catches up on coding, it&#x27;ll regress to the mean in other areas.
    • Damjanski7 hours ago
      thats so helpful - tysm
  • dzonga1 minute ago
    the last image - it&#x27;s barely visible to human eyes
  • weli10 hours ago
    Anecdotal, opinion:<p>Gpt is really good in vision stuff, or at least their MoE seems to be really cohesive. From my experience Claude models can be really good at language but the moment they need to look at a picture and decide why the design is not good what parts need improvement it degrades a lot. My easiest benchmark is giving them a screenshot of a feature in my app and tell it &quot;identify non-normative UI blocks and improve readability and consistency&quot;. Sol does a great job at re-structuring the page into composable units that build upon each other and the general looks and feels of the app. Claude tends to over-focus one one part while completely forgetting about the rest or the cohesion as a whole.
    • velcrovan10 hours ago
      Assessing the subjective quality of a thing is in my experience one of the worst ways to use any LLM.
      • TeMPOraL5 hours ago
        There&#x27;s a lot of objective principles and decisions that go into subjective quality; if you don&#x27;t know the field well, asking LLM for assessment is a good way to discover all that.
      • rib3ye10 hours ago
        anthropic frontend-design skill does a great job with it.
        • rafram10 hours ago
          Have you actually read the frontend design skill? It’s placebo at best. Very short and barely focused on design: <a href="https:&#x2F;&#x2F;github.com&#x2F;anthropics&#x2F;skills&#x2F;blob&#x2F;main&#x2F;skills&#x2F;frontend-design&#x2F;SKILL.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;anthropics&#x2F;skills&#x2F;blob&#x2F;main&#x2F;skills&#x2F;fronte...</a>
          • rib3ye9 hours ago
            Have you actually tried using it?
            • rafram8 hours ago
              Of course. It’s OK, but it tends to generate very cliched “AI” UIs with little originality. Despite the skill spending a lot of time coaching the model into avoiding that!
              • akoboldfrying28 minutes ago
                &gt; UIs with little originality<p>Sounds like the kind of UI I like. (Take me back to Windows XP...)
          • MallocVoidstar10 hours ago
            What an annoying time for GitHub to go down.
        • DaiPlusPlus10 hours ago
          My exposure to Claude-produced UIs is limited, but I have started to notice certain design trends they tend to have in-common, which might be becoming hallmarks of AI-produced UIs - the same way we&#x27;ve started noticing the clichés of low-effort LLM-generated text.<p>FWIW, the summary-description[1] of &quot;frontend-design&quot;[2] gives me a few things to pick at:<p>&gt; create polished code<p>Methinks only if you&#x27;re using it with a very popular framework like React. What happens if you ask Claude to make the UI in WinForms or MFC?<p>&gt; high-impact animations<p>That&#x27;s bad UX 101 right there: animations in a UI exist as an affordance to the user, and never for its own sake (e.g. macOS&#x27;s &quot;genie&quot; animation when you minimize a window to the dock exists so the user knows where they can restore the window from). The only people who actually want &quot;high impact animations&quot; in software are salespeople who want something for demo purposes.<p>&gt; generic system fonts, predictable purple gradients, and cookie-cutter components.<p>This screams wanting to be different for the sake of standing-out, not because it results in a better software product; users benefit when their software fits-in with platform conventions: if you refuse to use a stock checkbox &lt;input&gt; or &lt;select&gt; drop-down and instead use your own entirely custom component solely for aesthetic reasons then you are producing worse software. There&#x27;s nothing wrong with system-fonts, but your site will look ugly after your third-party font-host CDN shuts-down and turns into a walking CSRF factory.<p>&gt; thoughtful typography with unexpected font pairings<p>The above fragment set my alarm-bells off. Yikes.<p>&gt; scroll-triggered interactions<p>Not every web-page should be an Apple.com product brochure page. This is also a fantastic way to make your webpage horribly inaccessible.<p>------<p>The SKILL.md itself[3] grinds my gears too:<p>&gt; Approach this as the design lead at a small studio known for giving every client a visual identity that could not be mistaken for anyone else&#x27;s.<p>Claude has no way of knowing what designs are actually unique or not...<p>&gt; For web designs, the hero is a thesis. Open with the most characteristic thing in the subject&#x27;s world, in whatever form makes sense for it: a headline, an image, an animation, a live demo, an interactive moment<p>...this is <i>exactly</i> what everyone else&#x27;s web-pages look like!<p>&gt; For calibration: AI-generated design right now clusters around three looks: (1) a warm cream background (near #F4F1EA) with a high-contrast serif display and a terracotta accent; (2) a near-black background with a single bright acid-green or vermilion accent; (3) a broadsheet-style layout with hairline rules, zero border-radius, and dense newspaper-like columns<p>...I called this out weeks ago[4], lol.<p>and I could go on. This is all quite painful to read.<p>------<p>[1] <a href="https:&#x2F;&#x2F;claude.com&#x2F;plugins&#x2F;frontend-design" rel="nofollow">https:&#x2F;&#x2F;claude.com&#x2F;plugins&#x2F;frontend-design</a><p>[2] <a href="https:&#x2F;&#x2F;github.com&#x2F;anthropics&#x2F;claude-plugins-official&#x2F;tree&#x2F;main&#x2F;plugins&#x2F;frontend-design" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;anthropics&#x2F;claude-plugins-official&#x2F;tree&#x2F;m...</a><p>[3] <a href="https:&#x2F;&#x2F;github.com&#x2F;anthropics&#x2F;claude-plugins-official&#x2F;blob&#x2F;2a5cd1f39f0d8e5fbb68d77a13884f69c4b0c516&#x2F;plugins&#x2F;frontend-design&#x2F;skills&#x2F;frontend-design&#x2F;SKILL.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;anthropics&#x2F;claude-plugins-official&#x2F;blob&#x2F;2...</a><p>[4] <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49187385">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49187385</a>
      • keeganpoppen4 hours ago
        i&#x27;d say this is something that has gotten orders of magnitude better with recent releases than it used to be, fwiw
    • SkalskiP7 hours ago
      Hi! I’m the author of this blog. GPT-5.6 is much better at vision than previous GPT versions, but it’s still much weaker than Gemini 3.5 Flash or Gemini 3.7 Flash, which was released last week. One interesting approach is to use Gemini through a tool call.
      • Tactical453 hours ago
        This response is not relevant to the this comment
    • DaiPlusPlus10 hours ago
      What is a &quot;non-normative UI block&quot;?
      • weli10 hours ago
        Segments of the UI that don&#x27;t conform to any other existing established design or conventions
      • lelandfe10 hours ago
        areas that look weird
  • evrimoztamur10 hours ago
    Penny sample shown looks like failed EXIF orientation registered by the model&#x2F;harness. The coins are correctly marked, it&#x27;s rotated 90 degrees.
    • SkalskiP7 hours ago
      Hi! I’m the author of this blog. I had the same intuition, but together with the OpenAI team we figured out that the issue was image resolution. GPT-5.6 doesn’t handle large images well.
      • evrimoztamur2 hours ago
        OpenAI team sounds like they&#x27;ve misidentified the root cause for this particular case then.
  • bearjaws9 hours ago
    It is funny to me seeing Sol used for what a &quot;traditional&quot; AI model can do already (counting pills).<p>We have vision models for our pharmacy and I could never imagine taking the latency hit to use a Sol in our robotics, it would be likely 25-50x slower.
    • SkalskiP7 hours ago
      Hi! I’m the author of this blog.<p>I’m evaluating these VLMs to figure out which ones are good enough to auto-annotate my data, so I can fine-tune my detector.<p>I wrote a bit more about this here: <a href="https:&#x2F;&#x2F;x.com&#x2F;skalskip92&#x2F;status&#x2F;2080334344061694429?s=20" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;skalskip92&#x2F;status&#x2F;2080334344061694429?s=20</a>
      • mhaberl7 hours ago
        Did you evaluate any that could be self-hosted (or at least ow models), if so which one is the best you seen?
        • SkalskiP5 hours ago
          Take a look here: <a href="https:&#x2F;&#x2F;playground.roboflow.com&#x2F;evals" rel="nofollow">https:&#x2F;&#x2F;playground.roboflow.com&#x2F;evals</a>. We have few ~30B.
          • mhaberl4 hours ago
            Thank you!<p>It seems Qwen is kicking ass, and Fable made me laugh when I saw it all alone on the far right of the graph :))
    • kooi6 hours ago
      Agreed, this like asking a chainsaw to carve a wooden spoon. Impressive it can, but definitely not the right tech to scale.<p>LLM needs to setup an image classifier to use as a tool call.
      • bonoboTP4 hours ago
        Building a dataset is expensive, manual annotation is expensive. Datasets don&#x27;t exist in every niche.<p>I remember around 2013-15 people were scoffing at uses of deep learning CNNs for various things, because why don&#x27;t you just use an SVM on HOG features? Or face detection is solved, just use Viola-Jones.<p>What if you give the benefit of doubt and assume the author knows about alternatives and uses VLMs for their strengths? They use it to auto-annotate training data for regular deep learning models.
      • ramblerman6 hours ago
        Now maybe, but the gap is closing.
    • repeekad9 hours ago
      How are we supposed to pay off all these data centers and chips if you’re not willing to burn a microwave burrito worth of electricity for each prescription? Think of the benchmarks
  • fpgaminer9 hours ago
    Gemini 3 Flash should really be included in this comparison. Or at least 3.7. In most of my testing, 3.5 and 3.6 were both a downgrade in terms of vision capabilities, relative to 3, and at a much higher cost. 3.7 is slightly better than 3, finally.
    • bastawhiz5 hours ago
      3 Flash never left &quot;preview&quot; status and is listed as deprecated.<p><a href="https:&#x2F;&#x2F;ai.google.dev&#x2F;gemini-api&#x2F;docs&#x2F;deprecations" rel="nofollow">https:&#x2F;&#x2F;ai.google.dev&#x2F;gemini-api&#x2F;docs&#x2F;deprecations</a>
    • zuzululu2 hours ago
      but 3.7 flash is expensive for img inputs no ?
      • fpgaminer1 hour ago
        As usual for something so simple, Google&#x27;s docs seem unclear: <a href="https:&#x2F;&#x2F;ai.google.dev&#x2F;gemini-api&#x2F;docs&#x2F;pricing" rel="nofollow">https:&#x2F;&#x2F;ai.google.dev&#x2F;gemini-api&#x2F;docs&#x2F;pricing</a><p>For 3, pricing for image tokens was the same as text tokens. Since they don&#x27;t indicate a difference on 3.7, I would assume the same holds. And as far as I know the number of image tokens is the same for both (depending on the detail level you pick, but it&#x27;s generally around 1k per image).<p>So they&#x27;re about the same, 3.7 is slightly more expensive. At least until the end of the year (when they raise 3.7&#x27;s pricing).<p>Anyway, my point was that 3.5 tended to have worse performance and significantly higher costs. 3 and 3.7 are both better and cheaper than 3.5.
  • dllu6 hours ago
    Vision is still embarrassingly bad.<p>ChatGPT Pro with GPT 5.6-sol: <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6a834217-ca8c-83e8-a8e8-45d5b8797b67" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6a834217-ca8c-83e8-a8e8-45d5b8797b...</a><p>The puzzle: <a href="https:&#x2F;&#x2F;activityvillage-files.s3.eu-west-2.amazonaws.com&#x2F;s3fs-public&#x2F;images&#x2F;christmas_present_match_up_460.jpg" rel="nofollow">https:&#x2F;&#x2F;activityvillage-files.s3.eu-west-2.amazonaws.com&#x2F;s3f...</a>
    • TeMPOraL5 hours ago
      The second answer is far more revealing than the first:<p>OP:<p>&gt; <i>do you think you did a good job there</i><p>ChatGPT:<p>&gt; <i>I spent 15 minutes, emitted several fake-sounding “tracing the puzzle” progress updates, and then gave a confident permutation without showing that I had actually followed the lines correctly. It reads much more like I guessed than solved it. The only part I did well was obeying the “no Python or tools” instruction.</i><p>My observations:<p>1) Sarcastic tone suggests pre-prompting, or frequent (and therefore stored in memories) denigration of the model in past conversations. I&#x27;m leaning the former - it sounds like it was instructed to read admission of defeat.<p>2) The part about &quot;no Python or tools&quot; is setting the model up for failure.<p>I mean, this task is, for a human, basically a game of &quot;simulate a line following robot in your head&quot;. Pretty sure a VLM could solve that if it was allowed to do the same thing. Off the top of my head, an algorithm like:<p>1. Identify start and end points<p>2. Foreach start point, follow next pixel minimizing angle, until endpoint is reached.<p>3. Report answer<p>It&#x27;s literally what every human facing this task does.<p>EDIT:<p>My attempt - same image, prompt altered to allow for code (but still no search&#x2F;external checks), solved in 1&#x2F;5th of the time, correctly, and (going by thinking trace summaries that I don&#x27;t think show up in shared chats), basically the same way I&#x27;d approach it, by tracing the lines, coloring them as it goes.<p><a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6a834f76-8240-83ed-acff-0c67af399d49" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;share&#x2F;6a834f76-8240-83ed-acff-0c67af399d...</a><p>INB4: I know this is now not a pure vision check, but it really doesn&#x27;t make much sense to diss models for failing to solve tasks explicitly designed to teach humans to externalize computation that&#x27;s hard to do in their heads (i.e. kids, crayons, coloring paths).<p>Still, if such things are becoming a benchmark for tool-less evaluation, it&#x27;s only a matter of time until the models learn - much like humans learn in school - to follow algorithms mentally, essentially emulating an ad-hoc computer in their head.
      • dllu5 hours ago
        No pre-prompting, although I can&#x27;t be sure it didn&#x27;t use memories. &quot;No tools&quot; should theoretically have prevented it from looking up memories. FWIW, Grok and Gemini both failed in a similar way.<p>With Python, it was able to successfully solve it in 9 minutes: <a href="https:&#x2F;&#x2F;chatgpt.com&#x2F;s&#x2F;t_6a8350ecddfc81919328caf68de74861" rel="nofollow">https:&#x2F;&#x2F;chatgpt.com&#x2F;s&#x2F;t_6a8350ecddfc81919328caf68de74861</a><p>The real pain point is that at work, I use Codex and I&#x27;m currently working on a project that involves debugging some polyline topology, very similar to the path following puzzle. The vision is completely useless here.<p>Your VLM idea sounds good. Theoretically, the inverse problem (generating an SVG of a pelican riding a bike) can also be solved with a VLM that plans out how to draw it, not unlike a human planning out a path for their hand to follow.
  • mv410 hours ago
    Ironically, the pill counting example selected to showcase &quot;the best vision model&quot; can be easily solved with OpenCV template matching, a technology created 25 years ago.
    • maxime_cb9 hours ago
      I&#x27;m assuming you mean that this tech became available in OpenCV 25 years ago, but as it turns out, the underlying tech can be traced back much further, at least as far as 1977! :)<p><a href="https:&#x2F;&#x2F;ieeexplore.ieee.org&#x2F;document&#x2F;1674847" rel="nofollow">https:&#x2F;&#x2F;ieeexplore.ieee.org&#x2F;document&#x2F;1674847</a> G. J. Vanderbrug and A. Rosenfeld, “Two-Stage Template Matching,” IEEE Transactions on Computers, Vol. C-26, No. 4, pp. 384–393, April 1977. DOI: 10.1109&#x2F;TC.1977.1674847
      • mv48 hours ago
        Exactly my point. Template rotation is a trivial operation as well.
    • lebek9 hours ago
      The point is that it&#x27;s general. It can do this task and many other tasks and it doesn&#x27;t need custom development like OpenCV does. Of course if you only want to count pills and you want it to be cheap&#x2F;fast you&#x27;re still better off using OpenCV.
    • dekhn8 hours ago
      Basic Template matching has severe limitations around scaling, rotation, and perspective. In my experience it greatly underperforms compared to deep network object detectors. My experience- and I imagine others have different experiences- is that SIFT techniques also fail pretty badly with noisy data.
      • mv48 hours ago
        That&#x27;s correct, and I was specifically referring to the example chosen - where scale and perspective are known. Template rotation is relatively easy as well - but partial obstructions would pose a problem.<p>Another application where template matching would work brilliantly? Car counting in parking lots using satellite imagery.<p>Source: I did this [1] using OpenCV and template matching. Outperformed &quot;Cars Overhead with Context&quot; models.<p><a href="https:&#x2F;&#x2F;abcnews.com&#x2F;International&#x2F;satellite-data-suggests-coronavirus-hit-china-earlier-researchers&#x2F;story?id=71123270" rel="nofollow">https:&#x2F;&#x2F;abcnews.com&#x2F;International&#x2F;satellite-data-suggests-co...</a>
    • geysersam8 hours ago
      I&#x27;m sure a typical frontier model would also be happy to write that opencv script for you, and it would do it well.<p>That is certainly pretty far from what was possible 25 years ago.
      • mv48 hours ago
        It 5..10 lines of code. :)
  • kzrdude10 hours ago
    In the third vision bench result, Sol is 100% correct but the expected has 1 error. Seems like an oversight.<p>In the next bench, Sol looks like it’s correct again but the bboxes are rotated 90 degrees for some reason.
    • defrim9 hours ago
      Seems to be due to the detection area being not fully accurate. Green vs red shows the difference between actual and detected
      • kzrdude7 hours ago
        There is an extra green square where no egg is present, so it&#x27;s a false positive in the expected.
    • SkalskiP7 hours ago
      Hi! I’m the author of this blog and benchmark. You’re right. I’ll fix it in the ground-truth dataset. Thanks for pointing it out.
      • kzrdude7 hours ago
        Great, happy that it was helpful
  • schopra9098 hours ago
    From our experiments it’s the best video captioning model in the world by a mile. This was not the case a year ago.<p>When reasoning got introduced a year ago to GPT 5, on average the model performed worse than GPT4-o for short video clip captioning (Ie hallucinating actions that didn’t happen). The old GPT 5 was extremely finicky in terms of fps sample rate.<p>The other SOTA LLMs (like Gemini Pro) have clearly been optimized for long video understanding, since they can’t see almost anything sub-second (even if you up the frame sampling rate).<p>Sol is the first model we’ve seen to accurately caption complex sub-second movements (eg woman suddenly turns heard head to right). It’s robust to different fps sample rates so I can only guess that they trained on videos sampled at different fps.
  • faxmeyourcode8 hours ago
    It&#x27;s not clear to me from the article, are they asking sol to output bounding box coordinates with some kind of structured outputs?<p>Anecdotal but I&#x27;ve seen it use python to crop, zoom, and &quot;enhance&quot; (fiddle with sharpness and brightness) images to read sections of handwritten census data from the 1800s. Feels like that there might just be a mismatch of capabilities when it comes to straight outputting coordinates but I bet the model is better at actually finding the answer given any tools available. Which I get is a bit of an apples and oranges situation.<p>I&#x27;ve also tried to use it to identify an old pair of glasses and it didn&#x27;t stand a chance, so I do think it&#x27;s not quite there yet when it comes to some vision tasks.
  • lwarfield7 hours ago
    I currently have fable organize a bunch of 5.6 sol agents when working on my personal projects. This makes me wonder if I should add something along the lines of &quot;For tasks that involve visual analysis, have gemini 3.7 look at images generated.&quot;<p>Overall I&#x27;ve been hooked on using agents from different companies for what they are best at (Thanks to Theo). Fable is expensive, but unmatched for planning and top level organization of other agents. Sol is fast, will persistantly go after goals (sometimes to its detriment), and does well with computer use.
  • ALLTaken8 hours ago
    I actually favor Qwen3.8 and run it locally + use the Token-Plan on AlibabaCloud, when I need faster results. Kind of favor it over GPT5.6 Sol.<p>Also it seems to be more capable, need to test more, but I think it&#x27;s at least getting on par and it&#x27;s fully open-source and open-weights.<p>Here&#x27;s some benchmarks:<p><a href="https:&#x2F;&#x2F;benchlm.ai&#x2F;compare&#x2F;gpt-5-6-sol-vs-qwen3-8-max" rel="nofollow">https:&#x2F;&#x2F;benchlm.ai&#x2F;compare&#x2F;gpt-5-6-sol-vs-qwen3-8-max</a><p><a href="https:&#x2F;&#x2F;qwen.ai&#x2F;blog?id=qwen3.8#full-benchmark-table" rel="nofollow">https:&#x2F;&#x2F;qwen.ai&#x2F;blog?id=qwen3.8#full-benchmark-table</a> (incredible UI&#x2F;UX demos)<p><a href="https:&#x2F;&#x2F;venturebeat.com&#x2F;technology&#x2F;qwen3-8-max-arrives-with-a-bold-claim-it-outperforms-gpt-5-6-sol-max-and-fable-5-on-agentic-computer-use" rel="nofollow">https:&#x2F;&#x2F;venturebeat.com&#x2F;technology&#x2F;qwen3-8-max-arrives-with-...</a><p>EDIT: Am I early to the discussion, or is none else using Qwen3.8-max?
    • barrenko3 hours ago
      I thought Qwen 3.8 max doesn&#x27;t have vision?
    • ALLTaken5 hours ago
      huh, why am I being shadow banned?<p>Does YC have similar problems like those at wikipedia&#x2F;reddit? (wikipedia-editor-wars, or reddit-mod-wars)
  • apinstein4 hours ago
    It’s gotten so good that I now have infrastructure to render all mermaid&#x2F;plantuml in my project to png and have AI’s always load both text and image versions. And they are instructed to review the rendering as part of the diagramming cycle (for layout, salience, usefulness, etc). They can now produce useful diagrams that help reach shared architecture understanding.
  • kherud9 hours ago
    So far I haven&#x27;t seen a single model succeeding at transcribing sheet music, but I just tested it again with 5.6 Sol and it nailed the small test case. Fluently reading music requires multiple years of training for most people, but I feel like accurately following the horizontal lines trips up vision models in particular.
    • adrianh7 hours ago
      For a bespoke model that transcribes sheet music images well, check out our system at Soundslice: <a href="https:&#x2F;&#x2F;www.soundslice.com&#x2F;sheet-music-scanner&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.soundslice.com&#x2F;sheet-music-scanner&#x2F;</a><p>It&#x27;s not an LLM, it&#x27;s a custom thing we built. Here&#x27;s a comprehensive list of support for various notation glyphs: <a href="https:&#x2F;&#x2F;www.soundslice.com&#x2F;help&#x2F;en&#x2F;creating&#x2F;pdf-import&#x2F;294&#x2F;supported-notations&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.soundslice.com&#x2F;help&#x2F;en&#x2F;creating&#x2F;pdf-import&#x2F;294&#x2F;s...</a>
  • iamniels10 hours ago
    I understand why you would like to use an LLM for vision. I do it myself often enough. I don&#x27;t understand however, why the pill detection and counting is included in this benchmark. That is a task which you would perform with OpenCV right?<p>In my personal mini benchmark minicpm-v-4.6 scores amazingly well. Its a 0.8B model which runs fine on many consumer hardware.
    • throwup23810 hours ago
      Generating datasets to train more efficient models is a common use case for VLMs, especially frontier ones. It makes it much cheaper to create that initial dataset and you can abuse the nondeterminism of LLMs to identify data for human review (if they don’t converge, escalate to a human).
    • rhplus9 hours ago
      Especially the pill counting example. The best model was shown at 81.1% accuracy, which is a terrible rate for pharmacy scenarios. It seems like implementors would be better off instructing the models to use deterministic tools (like OpenCV) until the models are at 99.99% accuracy (or whatever an acceptable error rate is for pharmacy techs).
    • jacquesm8 hours ago
      I think that is because people perceive OpenCV as &#x27;hard to use&#x27; and LLMs as easy to use.
      • TeMPOraL4 hours ago
        OpenCV is no longer hard to use, it just takes longer. Still, a little more complicated than asking LLM to count.<p>To use an LLM, you just prompt it with an image + text saying &quot;count the pills in this image&quot;.<p>To use OpenCV, ... you just prompt an LLM with an image + text saying &quot;count the pills in this image, using OpenCV instead of eyeballing it&quot;.<p>(I like to throw in &quot;produce intermediary artifacts so I can see the process&quot; for more difficult tasks; this helps the model avoiding making hallucination-prone leaps and gives more opportunities to self-correct. At a cost of extra time and tokens, of course.)<p>Using OpenCV without an LLM? Nah, not touching that, I don&#x27;t have free weekends to waste anymore.
        • jacquesm2 hours ago
          I no longer use it but never felt it was particularly complicated, but since the days of resnet there are <i>much</i> faster ways to the goal.
  • drak0n1c2 hours ago
    Seed Turbo 2.1 is incredibly detailed in describing every physical feature. I use that one for vision tool calls through Venice API.
  • chasd0010 hours ago
    One of my friends (and BIL) own an architecture firm. They use AI to generate and quickly update renderings but they run into the equivalent of the 6 fingered hand problem. I sent him this article I wonder if the updated models can catch and fix mistakes made by previous models.
    • Mashimo9 hours ago
      This article is about vision, not image output.
      • stavros9 hours ago
        Hence the &quot;catch and fix mistakes&quot; part.
  • jug9 hours ago
    I really like the combo 5.6 Luna &amp; Sol for price and performance and would be perfectly happy if they stayed here for a moment without mucking about with sidegrades that I think AI evolution has often felt like lately.
  • ParanoidShroom9 hours ago
    I run the free service <a href="https:&#x2F;&#x2F;countrx.app&#x2F;" rel="nofollow">https:&#x2F;&#x2F;countrx.app&#x2F;</a> so i have some idea what goes into counting.<p>The performance as a general model is indeed really impressive and i think they might actually win compared to fine tuned models.<p>Their feedback loop of training on user data is incredibly strong. I&#x27;ve learned that lots of accuracy results depends on threshold configs, which llms should be able to dynamically set.<p>Or the future will develop in llms using fine-tuned models as tools? Inference cost and speed does still seem to be below user expectations.<p>But for being able to one shot with this accuracy... IMPRESSIVE
    • IncreasePosts8 hours ago
      How are you running it for free? Are you self funding or do you have sponsors?
      • ParanoidShroom5 hours ago
        Self funded. It&#x27;s a custom trained efficient model on CPU so it&#x27;s borderline free
  • bob102910 hours ago
    I&#x27;ve decided it&#x27;s &quot;good enough&quot; after I saw it properly quote a string of text that was very roughly highlighted within a nested visual context. It also identified the context correctly (modal inside webapp inside screenshot of user desktop).
  • 5555watch10 hours ago
    All of your use cases are very advanced.<p>I recently used it at grocery stores in a foreign country. Photographed the whole aisle and told it to find Y (detergent, softener, glue, sour cream, whatever), at the same time recommend the best Y for whatever reason. Worked marvelously, including the cases where the object wasn&#x27;t present and it told me there was nothing useful.<p>I asked then, can you crop the exact image of how does the item look like and where is it in the aisle - did that perfectly as well.<p>I will add that all frontier models were fine with such tasks from the early 2024&#x27;s.
  • dangoodmanUT1 hour ago
    I hate these &quot;The best X thing Y has ever released&quot;.<p>Unlike when Apple says &quot;it&#x27;s the best iphone we&#x27;ve ever made&quot;, LLMs are more or less interchangeable. So &quot;OpenAI&#x27;s best model&quot; means nothing if &quot;Anthropic wipes the floor with them&quot; or &quot;[open weights model] is 10x cheaper for 1% less quality&quot;.<p>As a reader, it feels like these titles are click bait.
  • slibhb6 hours ago
    One of the use cases I&#x27;ve wondered about for AI is giving it a picture of the &quot;spice wall&quot; in a grocery store and asking it to find all jars of e.g. cardamom. This takes me an annoyingly long time to do when I&#x27;m shopping, so it would actually be useful.
  • 1saadcodes7 hours ago
    5.6 Sol looks nice, but the Gemini 3.5 Flash comparison is interesting. It’s cheaper and still came out ahead on detection and counting, which doesn&#x27;t really give me much of a reason to use Sol since Flash is much cheaper and hence much easier to scale. Not to mention we now have 3.6 Flash too
    • ComputerGuru5 hours ago
      We have 3.7 Flash now, actually, and it costs just a hair over the old 3 Flash Preview while being better!
  • prathje10 hours ago
    I would love more vision benchmarks! Once I asked the model to inspect a completely black picture and it hallucinated a nice wooden kitchen wall. Took me some time to figure out where the kitchen came from...<p>I usually go to <a href="https:&#x2F;&#x2F;arena.ai&#x2F;leaderboard&#x2F;vision&#x2F;pareto" rel="nofollow">https:&#x2F;&#x2F;arena.ai&#x2F;leaderboard&#x2F;vision&#x2F;pareto</a> for a nice overview of current models.
  • WarmWash9 hours ago
    It&#x27;s vision capabilities poisoned my cucumber bed, misidentifying the malaise and having me spray them down with water, which only spread the fungus that gemini later informed me was actual cause, which I went and checked myself.<p>I hope that whatever was lost at GDM in the last few months, didn&#x27;t include their extra focus on vision capabilities.
  • sscaryterry10 hours ago
    My anecdotal evidence says its still as blind as any other model, it has no taste, no attention to any sort of detail.
    • howdareme10 hours ago
      How can a vision model have taste?
      • sarreph10 hours ago
        If you&#x27;re doing any kind of inference that is multi-modal and non-factual, opinions and biases will affect any kind of assessment of a visual that you provide to a model.<p>For example, a UI &#x2F; UX professional being asked to appraise a website screenshot may determine that the image in question has &quot;desirable&quot; traits which are inherently not deterministically measurable. Such as, if the interface elements have strong information hierarchy, or if they are deemed to be &quot;fashionable&quot; with current UI trends.
        • DaiPlusPlus10 hours ago
          &gt; if the interface elements have strong information hierarchy<p>...but that&#x27;s an example of a UX&#x2F;usability matter that can be assessed objectively and non-subjectively.
          • sarreph9 hours ago
            I disagree.<p>Is 16 px or 14 px a better font-size value for a subheading, in a hypothetical layout? Immediately that kind of decision, where both options are <i>objectively good</i> for 12 px paragraph text, suddenly becomes an issue of <i>taste</i> that cannot be evaluated crudely by an algorithm.
      • sscaryterry10 hours ago
        Replace taste with consistent if that helps you. Can it follow a design system...
        • yreg10 hours ago
          As a design system engineer I usually have to fight against the taste of the designers. (And I consider it natural.)<p>But, if you have a proper well documented design system and you tell the LLM to use the DS and to avoid styling hacks they can generally do it. Even the dumber ones than Sol 5.6.<p>Of course only if the design is achievable in the design system.
          • sscaryterry10 hours ago
            This is not my experience at all.
        • velcrovan10 hours ago
          So, formulaic output…the opposite of taste
          • sscaryterry10 hours ago
            Not really. Compliance with the letter of the law doesn&#x27;t mean the intent is complied with.
  • cdolan9 hours ago
    Luna is pretty strong as well. been using it for projects the last two weeks and its strong
  • trumbitta210 hours ago
    &quot;Best iPhone ever&quot; vibes.
  • wahid_seddiqi7 hours ago
    Do you think we’re getting closer to models that actually understand what they’re seeing, or are they just getting really good at recognizing patterns?
    • Culonavirus7 hours ago
      All I&#x27;m fine with for now is that I can almost exclusively communicate with Sol through collages and my scribblings (all kinds of web page &#x2F; block screens with all kinds of arrows and text all over the place) This was not practically ppossible in 5.5 and a tragedy in 5.4. Not sure how much weight is codex uploading in higher res carrying here but it&#x27;s great to work with.
  • adroitboss10 hours ago
    I didn&#x27;t expect Gemini 3.5 Flash to top basically every metric in this article.
    • SweetSoftPillow10 hours ago
      In my practice Gemini models are far better than anything on the market in terms of vision, also it&#x27;s worth to mention that current Gemini flash is 3.7, so it got 2 updates since 3.5 which beat GPT-5.6 Sol in this comparison.
    • SkalskiP7 hours ago
      Hi! I’m the author of this blog. I wrote it 4 weeks ago, and it’s already a bit outdated. Gemini 3.7 Flash came out last week, and considering the price, it’s easily the best vision model right now: <a href="https:&#x2F;&#x2F;x.com&#x2F;skalskip92&#x2F;status&#x2F;2088032652301304121?s=20" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;skalskip92&#x2F;status&#x2F;2088032652301304121?s=20</a>
    • LollipopYakuza10 hours ago
      Same. I scrolled back up to see if I read the title correctly. It&#x27;s important to note that it is the best... OpenAI released. Not the best overall.
    • WarmWash9 hours ago
      Gemini has long been the vision champion, but there aren&#x27;t many benchmarks and coding is where all the hype is.<p>Demis had a pretty big interest in vision, more so than text, so I hope they don&#x27;t lose that with all the recent shuffling.
  • sam0x171 hour ago
    &gt; GPT 5.6 Sol is the best &quot;vision&quot; model OpenAI ever released<p>I mean I should hope so, as it is also the latest one
  • criddell9 hours ago
    Are any of these vision benchmarks binocular in order to introduce depth perception?<p>I keep waiting for these AI companies to assemble the parts into a great autonomous driving module.
  • comboy10 hours ago
    Does any popular NVR make a good use of LLMs (especially local models) getting decent at vision?
    • eks3918 hours ago
      I&#x27;ve been using Reolink for years and been very satisfied with it.<p>The only quip is the default UI isn&#x27;t very good. When changing that reaches the top of my priority list, I&#x27;ll switch it since they don&#x27;t force you into a walled garden. Plan is to run it through frigate into HomeAssistant and use a UI from them. I&#x27;ve never used frigate before though so it&#x27;ll be a learning process if plug and play solutions aren&#x27;t already available
  • logicallee10 hours ago
    I agree. It did <i>very</i> well on an extremely challenging task.<p>I asked it to recognize and draw the very faint reflection of what I was wearing, visible in only a tiny black part of a very brightly lit poster behind glass.<p>In addition, the poster itself also happened to contain similar clothing.<p>You can see the reference images and its output in my writeup here: <a href="https:&#x2F;&#x2F;medium.com&#x2F;@rviragh&#x2F;gpt-5-6-sol-very-good-image-recognition-and-generation-5b0d0329a46f" rel="nofollow">https:&#x2F;&#x2F;medium.com&#x2F;@rviragh&#x2F;gpt-5-6-sol-very-good-image-reco...</a><p>While a human can focus on the reflection easily, this is an enormous challenge for a vision model. It&#x27;s very impressive.
  • Razengan10 hours ago
    For the last 2 weeks I&#x27;ve been trying to get Codex to &quot;outpaint&quot; a wonderful image it generated as placeholder art for a level background.<p>After I increased the game&#x27;s resolution, I asked it to increase the image&#x27;s size while keeping the same scale and existing content, and gosh, it constantly keeps getting something wrong no matter what I tell it, even on Sol Max with the $100 Pro subscription.<p>An organically-grown meat-based pixel-artist could have recreated the image and more within 2-3 days, in exchange for food and shelter.
    • dev_hugepages10 hours ago
      I&#x27;m unsure why you&#x27;re using an LLM to generate images. Don&#x27;t we already have models (some made by the same company) that do this?
    • sscaryterry10 hours ago
      &gt; it constantly keeps getting something wrong no matter what I tell it<p>This 100%
    • thatcat10 hours ago
      did you try segmenting it first?
      • Razengan9 hours ago
        At first I intended to create a tileset and asked it for several variations of what a hypothetical tilemap created from the planned tileset would look like.<p>The previews it generated were amazing but wouldn&#x27;t really be possible as a grid-based tilemap, with lots of clusters and overlaps of elements of varying sizes.<p>So I just decided to use the preview as a static scrolling background, but it&#x27;s been a pain to get it to add more content around the edges that still tiles with the existing image at the same scale.
  • slybot3 hours ago
    Am I the only one who cannot read the date on the blister pack even fully zoom in my phone?<p>If that is the full quality image given to the model, I think it&#x27;s not surprising that the model confused with 03&#x2F;2022.
  • terhechte7 hours ago
    Fuck ack. I&#x27;m working on a new benchmark that combines strong visual requirements with tool and coding requirements. I haven&#x27;t even tested Sol yet, but between Sonnet, Terra &amp; Luna I already see much better results from OpenAI&#x27;s models. I&#x27;m not releasing anything yet as I still have issues in my harness that need to be fixed.
  • RugnirViking9 hours ago
    It&#x27;s really quite good! I was amazed recently by its utter inability to read some faded handwritten cyrillic on the back of a wood carving - 3 or 4 words only, reasonably clear letter forms I found recently, and then stepped back a bit and thought about how insane that was as a benchmark - I just expect it to work so reliably on other OCR and translation tasks that it was surprising to encounter such a failure
  • iamleppert10 hours ago
    Where are the Qwen benchmarks in this? I would be more interesting to see how Qwen performs.
    • SkalskiP7 hours ago
      Hi! I’m the author of this blog. I regularly benchmark new VLM releases. You can check the results for Qwen3.8-Max and Qwen3.8-27B here: <a href="https:&#x2F;&#x2F;playground.roboflow.com&#x2F;evals" rel="nofollow">https:&#x2F;&#x2F;playground.roboflow.com&#x2F;evals</a>
    • ImageXav9 hours ago
      Me too. This is an interesting comparison but in my experience Qwen and Gemini have typically been the top contenders for image related tasks. For that reason it would be great to have the comparison here, as I&#x27;m not surprised by Gemini&#x27;s dominance over the other models.
  • fooker8 hours ago
    I&#x27;m a little bit disappointed that vision seems to fall before language at scale.<p>It seems pretty counter intuitive that we can&#x27;t do vision significantly better with specialized techniques.
  • TZubiri8 hours ago
    Which is to say, still not ready for any production workloads yet. As in, it cannot reliably count the amount of objects in an image.<p>Still very impressive, but nowhere near the text chat revolution. OpenAI still trying to strike their second lightning
    • chistev1 hour ago
      <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=46444508">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=46444508</a>
  • catigula10 hours ago
    Still not quite as good as gemini.
  • fintuner9 hours ago
    [flagged]
  • johnxianren7 hours ago
    [dead]
  • alessandrobinda8 hours ago
    [dead]
  • zyvop110 hours ago
    [flagged]
  • hathym10 hours ago
    [dead]
  • ZeroDayDreamer7 hours ago
    [dead]
  • hn7jmxa7oc10 hours ago
    [dead]
  • CurbStomper9 hours ago
    [dead]