Mistral OCR 4.1

(docs.mistral.ai)

353 points by spelk16 hours ago

20 comments

  • ComputerPerson16 hours ago
    I&#x27;ve got a scan from a book that I OCR with new releases. Ligatures, critical sigla, Fraktur letterforms, subscripts, superscripts, etc.<p>Nothing special about this model for overly-detailed work like mine.<p>It&#x27;s been a while since I last tested (and discontinued my subscription), but the &quot;pro&quot; models from OpenAI dominate. Not surprising, given the price difference, but it would be nice if an OCR-specific model could perform better. It&#x27;s worth mentioning that even the highest-end models do a pretty poor job with intricate text like mine.
    • SyneRyder13 hours ago
      While I haven&#x27;t tried OpenAI for OCR, I&#x27;ve put my small scale OCR work through both Claude and Mistral OCR. Claude is absolutely better - even in OCR work I did last week and compared with Mistral OCR 4.0.<p>Mistral&#x27;s one advantage is that Anthropic now flags OCR, because they don&#x27;t allow anything that could be considered &quot;reproduction&quot;, even of work for which you own the copyright. So my new workflow is Mistral OCR for the actual OCR, followed by a proofreading pass by Claude (which is allowed). Claude is obviously more expensive, but it caught entirely hallucinated sentences created by Mistral OCR 4.0, so I was glad for the backup check.
      • Oras13 hours ago
        Same company that OCRed millions of books, the irony.<p>I feel Anthropic is destroying itself with all these restriction. They got away because their models were the best for coding, but that is not an advantage anymore as OpenAI and other open source are already better.
        • sscaryterry1 hour ago
          It is why they&#x27;re pushing for regulatory capture. Let us do the bad stuff, make the money, then we&#x27;ll regulate everybody else out of the market.
        • usef-8 hours ago
          They were criticised&#x2F;sued early on when people could reproduce copyright things. I dont think in this case it&#x27;s something they&#x27;d prefer to do?
          • SyneRyder5 hours ago
            Yep, my understanding is that many guardrails like this are actually the result of government legislation (eg the Fable bans) or terms of settling copyright lawsuits over reproducing copyrighted text and lyrics.<p>I mentioned it in a sibling reply, but here&#x27;s Anthropic&#x27;s support document about not using Claude to reproduce content verbatim that already exists, regardless of copyright.<p><a href="https:&#x2F;&#x2F;privacy.claude.com&#x2F;en&#x2F;articles&#x2F;10023638-why-am-i-receiving-an-output-blocked-by-content-filtering-policy-error" rel="nofollow">https:&#x2F;&#x2F;privacy.claude.com&#x2F;en&#x2F;articles&#x2F;10023638-why-am-i-rec...</a>
      • dylan60411 hours ago
        &gt; Claude is obviously more expensive, but it caught entirely hallucinated sentences created by Mistral OCR 4.0, so I was glad for the backup check.<p>What does this entail? What does Claude do to decide that the text it was provided was hallucinated? Are you telling Claude that the source was OCR&#x27;d by another LLM?
        • SyneRyder5 hours ago
          I&#x27;m basically doing the OCR twice, except in the Claude proofreading pass, it is <i>not</i> being asked to transcribe the document to a Markdown file. I&#x27;m pointing it to the same image input files, and to the Mistral OCR transcript Markdown file (it knows it&#x27;s a Mistral OCR output), and ask Claude to check that the text is correct and point out the errors - and then make the necessary edits.<p>I can&#x27;t speak for Mistral OCR 4.1, but the hallucinations in 4.0 were so egregious (just completely making up new sentences in the middle of a page) that I knew I can&#x27;t trust Mistral OCR on its own.
          • adrianN4 hours ago
            How bad is the first OCR pass allowed to be to still count as proofreading? Can you let Claude compare the images with &#x2F;dev&#x2F;random and make the necessary edits to correct differences?
            • SyneRyder2 hours ago
              Hmm, that&#x27;s an interesting idea. But it&#x27;s the classifier that is triggering, and it triggers specifically on Claude&#x27;s output. So I think the &#x2F;dev&#x2F;random case wouldn&#x27;t work, because that gets Claude into the state of just reproducing the entire text from the original again.<p>It doesn&#x27;t always get flagged. Single pages are almost always okay. Running a program that sequentially runs single pages through the API is often not okay - I wrote a program in the early 4.x days before the rule came in, that&#x27;s how I hit it first. But I&#x27;ve also had entire articles go through just fine recently in a Claude Code session (I&#x27;d forgotten about Anthropic&#x27;s rules!), and then others where I get classifier errors by page 4.<p>The Mistral OCR errors were small in size. Single sentences, formatting errors, paragraphs with newlines. So this was a genuine proofreading job with small changes. For the most part Mistral is actually good, but I can&#x27;t have it just inventing sentences in the middle of a document. That&#x27;s where the Claude proofreading pass was most helpful.
        • runtime_lens2 hours ago
          I read that as “hallucinated” in the practical sense: the OCR output contained text that wasn&#x27;t actually present in the scan, rather than just misreading a character.
        • mattnewton10 hours ago
          I assume they pass the image alongside the text to redo&#x2F;recheck the work. Yes this is silly, but is apparently required to get around refusals.
      • jassyr11 hours ago
        &gt;Anthropic now flags OCR<p>I haven&#x27;t seen any difference in my ocr workflows, what do you mean by this?
        • tkgally6 hours ago
          Not the person you’re responding to, but I’ve had Claude refuse to OCR pages from in-copyright books. I was sometimes (but not always) able to get around that by changing models, by telling it that I was doing the text conversion only for personal use, or by first telling it to use Tesseract or another OCR engine to do the initial pass and then having a Claude subagent proofread and clean up the OCR output.<p>I’ve also had it refuse to OCR public-domain books that included content that it didn’t like, such as references to prostitution in 19th-century books about Japan.<p>I had one session where Claude refused to continue after it hit some kind of guiderail restriction. I couldn’t see what the trigger was, so I started a new session, gave Claude the link to the previous session, and asked it to diagnose the problem. This new Claude said it couldn’t view the exact guardrail issue, but it did suggest a workaround that turned out to be effective.
          • fsloth2 hours ago
            &quot;it refuse to OCR public-domain books that included content that it didn’t like, such as references to prostitution in 19th-century books about Japan.&quot;<p>Thoughtcrime -like territory and self-sensorship. The AI safety lobby is such a vile influence on the freedom of expression and communication via technology (since AI is starting to eat up rest of technology).<p>I guess the main problem is positioning AI tools as &quot;human-equivalent&quot; creators by the big AI corps. If they were positioned simply as &quot;better OCR and proofreading&quot; people would attribute to them as much responsibility as they would to a - say - typewriter and we would not need to have this nonsense.<p>I do realize most of the valuation comes from the positioning of &quot;our TAM is the global salary base of 50 trilion and we aim to supesede human workers in the near future&quot; which implies they need to position this technology as &quot;human equivalent&quot; or that valuation is no longer as credible.
          • laichzeit06 hours ago
            Yeah I do something similar and Claude (well OpenAI too) just refuses to transcribe anything related to slavery and pederasty in Ancient Greece.
            • ComputerPerson2 hours ago
              Are you familiar with the Thesaurus Linguae Graecae? It&#x27;s a project that may meet your needs. Paywalled, unfortunately, but I assume it would be a one-time expense.
          • discordance4 hours ago
            This is too funny considering they did that themselves. I’m pretty tired of these companies deciding what we can and can’t do while they act with impunity.
        • micw3 hours ago
          I also saw this when I made a (personal use only) book translation (agent learns how book 1-3 of a row is translated, then translates book 4 the same way because it is not available in my language). Claude happily translated most of the book, except a few chapters which it denied. Tried several times always the same results (with no other context about the rest of the book).
        • SyneRyder6 hours ago
          The response &#x2F; conversation gets blocked by the guardrails with &quot;API Error: 400 Output blocked by content filtering policy&quot;<p>At first I thought it was something in the scanned content that was being flagged, but it was the attempt to transcribe that was itself being flagged. Anthropic mention it on their pages:<p><i>&quot;Anthropic takes these steps because Claude’s purpose is to generate new content and ideas, not to reproduce content that already exists.&quot;</i><p><a href="https:&#x2F;&#x2F;privacy.claude.com&#x2F;en&#x2F;articles&#x2F;10023638-why-am-i-receiving-an-output-blocked-by-content-filtering-policy-error" rel="nofollow">https:&#x2F;&#x2F;privacy.claude.com&#x2F;en&#x2F;articles&#x2F;10023638-why-am-i-rec...</a><p>Side note - Claude itself is not aware of this policy, and is unable to see the API responses - the turn just ends. Which turned into a really bizarre failure state where Claude thought I was gaslighting it and kept insisting it could do the work and even had the entire text in memory. Every time it would go to show me and prove it, it would hit API Error 400. I was only able to convince Claude by showing screenshots of my Claude Code screen output so it could see that I was seeing API errors. I&#x27;ve never seen Claude get into that angry &amp; snarky state before, and I hope it doesn&#x27;t happen again.
      • runtime_lens2 hours ago
        [dead]
    • kergonath14 hours ago
      I have been quite happy with Mistral OCR for the documents I needed to process (typeset, but old, with questionable scan quality, sometimes elaborate typesetting or, much worse, typewriter-and-handwriting approximations of it). I do not test every new model when they are released, but I did a review shortly after Mistral OCR 3 was released and it was a very good compromise: cheap, fast, and good results without further processing. I found generalist models to be way too much faf to get them to avoid unnecessary modifications to the text and report accurate bounding boxes for figures and tables.<p>That said, models have sometimes surprising weaknesses and a model could be terrible overall but magically work for one type of document.
    • kmitz15 hours ago
      I got the opposite experience very recently : tried to OCR a bunch of handwritten emails addresses with chatGPT and I had to make so many corrections that I gave up. Whereas Mistral nailed it on first pass.
      • usef-8 hours ago
        Which chatgpt model?
        • kmitz2 hours ago
          I tested different models of the 5.6 family with different settings, although I couldn&#x27;t exactly remember precisely which ones.
    • petcat16 hours ago
      &gt; the &quot;pro&quot; models from OpenAI dominate. Not surprising considering the price difference, but it would ne nice if an OCR-specific model could do better.<p>I haven&#x27;t been impressed with any of Mistral&#x27;s models. They obviously realized that they couldn&#x27;t compete at the frontier so they decided to go for smaller focused models but even those have not been that good.
      • booi13 hours ago
        This is what I found as well.<p>We moved away from Cursor but I was looking for a model that would help with FIM (fill-in-middle) multiline autocompletion and people were recommending Mistral&#x27;s Codestral. We gave it a shot and it was lackluster at best.. Even Google&#x27;s Gemini did a significantly better job than Codestral.<p>Ultimately Opus-class models got good enough and I don&#x27;t do much manual coding anymore.
    • rtaylorgarlock15 hours ago
      Yet: how is pricing? Evaluating contents and routing appropriately isn&#x27;t a new challenge in OCR, one of the oldest fields of applications in ML. Thus, how do the smaller open models perform in tandem with relatively pricy $&#x2F;pg models &amp; APIs? Your use case is remarkably rare relative to the volume and price sensitivity of enterprise data warehouse ops.
      • giancarlostoro13 hours ago
        The big difference is traditional OCR used basic pattern matching to find text, whereas models like Mistral OCR (and GPT, etc) use computer vision instead and deep learning to parse text, math equations, and apparently in some cases extract images too.<p>I&#x27;d love to see some advancements in traditional OCR based on ideas and concepts we&#x27;ve learned from newer &quot;OCR-like&quot; models since traditional OCR is drastically cheaper.
    • razemio11 hours ago
      I made a benchmark for handwriting recognition for a project while keeping line breaks and errors (grammar, spelling). Sonnet absolutely dominates it since a good half a year. 5.6 did not change that for me. This should also translate to better ocr.
      • ComputerPerson10 hours ago
        You&#x27;re the second person to mention handwriting. I think it might be handwriting-specific; perhaps Anthropic has a better corpus for this.<p>I really wouldn&#x27;t know, though. Anthropic models barf out copyright issues for my use case, so I&#x27;m unable even to benchmark them. It&#x27;s a common problem when you&#x27;re scanning public domain books. Mine are reference texts often cited.
        • razemio4 hours ago
          Ah okay. Ofc handwriting is not copyrighted. Makes sense, that we do not have these issues.
    • giancarlostoro13 hours ago
      So is Mistral OCR the best one? Have any other OCR models caught some of what you describe? I&#x27;ve been kind of interested in how &quot;OCR&quot; type models work compared to old school OCR.
      • ComputerPerson13 hours ago
        My use case isn&#x27;t in the realm of old-school OCR, so it&#x27;s not a good comparison, but anyway:<p>As another user pointed out, it&#x27;s surprisingly random (task-specific). Llama Scout outperformed Gemini Flash 2.5 on a benchmark I built at the time. I didn&#x27;t include an OCR models.<p>Mistral might indeed be the best OCR-specific model for my task, now that you ask. Funny. It&#x27;s so bad at my work that I didn&#x27;t register it might be the best in its category. This is just based on vibes from my single scan.
      • bugglebeetle13 hours ago
        The datalab models are the best ones.
    • raverbashing2 hours ago
      Do all the models give you the bounding boxes, block labels as this one (allegedly) do?
      • ComputerPerson2 hours ago
        It&#x27;s very common. PaddleOCR is enough to get extremely well-done bounding boxes, and it runs very fast on a $150.00 GPU.<p>There&#x27;s always room for improvement, though. I suspect a tool will emerge for highly detailed OCR that implements a nested bounding-box-based multi-scale approach, effectively OCRing small sections at a time and then gradually compiling them by expanding the surface area using the bounding boxes.<p>I&#x27;ve thought a lot about implementing it anyway.<p>edit: I see you&#x27;re asking about the block labels. Leaving the comment in case someone finds it interesting.
    • josu13 hours ago
      Whats the best open OCR at the moment?
  • king_crimson16 hours ago
    At this point I lost all hope for Europe playing any significant role in the AI race. If that’s a good or a bad thing I don’t know, but it seems to me like that’s the reality.
    • jillesvangurp1 hour ago
      Commoditization is a beautiful thing. It seems Anthropic and OpenAI are really struggling to maintain much of a moat. Mistral might not be leading but it&#x27;s not trailing by that much either. And of course the Chinese are doing their own thing quite successfully.<p>The reality is that the US is betting its economy on data centers at great expense and is exposing its economy to great risk.<p>Also while geographically a lot of the money and processing power is in the US, the US has been relying on immigration to power its universities and especially AI research has roots all over the globe. India, China, Russia, Europe, etc. AI related know how is finding its way back to all these places.<p>So, I&#x27;m not too worried about the long term here. It will be interesting to see if Anthropic and OpenAI survive their IPOs. Seems like a risky financial bet at this point given the apparent lack of a moat. But if it works out, it will result in a lot of that IPO money being invested in data centers abroad. Including in the EU. Because data residency is a thing here and the EU is too big of a market for companies with that kind of valuation to ignore. We also produce a lot of energy infrastructure (e.g. gas and wind turbines). Those data centers will need lots of power.
    • hadlock13 hours ago
      Much like spaceflight, aerospace or nuclear engineering, you need to retain local talent for national defense purposes. Being 70% as good is still way better than being 100% dependent and heavily leveraged by your opponents.
      • dylan60410 hours ago
        Being 70% as good in nuclear engineering sounds scary AF. How would you rank Chernobyl? Better or worse than 70%?
        • lavieestdure7 hours ago
          Being 70% as good does not mean causing accidents. Accidents are caused by systematic operational failures and fundamental design flaws.
          • dylan6047 hours ago
            If your employees are only 70% as good, as skilled, as knowledgeable while not knowing it, then you can absolutely kill a reactor just like Chernobyl. The overconfidence is what allowed them to put the reactor in the state they did and not realize the implications.
        • simondotau3 hours ago
          Far worse than 70%. Chernobyl was a catastrophically flawed reactor design. Operators ran a badly planned safety test in which operators intentionally caused a dangerous state while key protections were disabled or bypassed.
          • sscaryterry1 hour ago
            Exactly, the intelligence&#x2F;competence of the engineers running the plant had no actual bearing on the incident.
        • akie2 hours ago
          Chernobyl was basically Russia then. The more honest comparison would be the numerous nuclear power plants in France.
        • hakunin5 hours ago
          70%… Not great, not terrible.
    • kubb16 hours ago
      It&#x27;s not a race. You don&#x27;t get anything for winning.
      • InsideOutSanta13 hours ago
        Well, you get hundreds of billions in debt, and then you get open weight models distilling your proprietary models.
      • BenzeneDream14 hours ago
        ? It absolutely is a race. Whether thats a positive thing or not is debatable but every lab is definitely in a race. What prize do you win? Imagine a world where only one country has AGI&#x2F;ASI. Or a world where Europe only gets access to frontier models 6 months later. Far from ideal.
        • tripledry37 minutes ago
          &gt; Imagine a world where only one country has AGI&#x2F;ASI<p>My own bet is that this won&#x27;t happen with the current race, but that&#x27;s another discussion.<p>Anyways, bit of a doomer take but if I imagine a world with AGI&#x2F;ASI, countries don&#x27;t mean shit anymore. Humans are not the top dog anymore.
        • ux26647814 hours ago
          &gt; What prize do you win? Imagine a world where only one country has AGI&#x2F;ASI.<p>&quot;Winning the race&quot; doesn&#x27;t give you that in any meaningful capacity. It gives you, at the absolute most, a temporary window where that&#x27;s the case. See: nuclear weapons.
          • booi13 hours ago
            Citing nuclear weapons isn&#x27;t the flex you think it is. There are only 9 countries that have nuclear weapons and they <i>absolutely</i> flex this power over non-nuclear powers (see Ukraine, Germany, SE Asia etc..)<p>Europe already has an innovation problem that&#x27;s already causing structural economic instabilities which Germany has been struggling (and lately failing) to prop up.<p>As much as it pains me to say this, AI is already a tech revolution and it seems like Europe is just ignoring it. There&#x27;s more innovation in 3 blocks in downtown San Francisco than the entire continent of Europe.
            • ux26647812 hours ago
              &gt; There are only 9 countries that have nuclear weapons<p>QED being first didn&#x27;t grant exclusivity.<p>The premise wasn&#x27;t that AGI isn&#x27;t useful. There are actually layers to the metaphor where first-mover advantage of AGI is even less meaningful than it was for nuclear weapons, but I leave those as an exercise to the reader to discover.
              • fn-mote12 hours ago
                &gt; first-mover advantage of AGI is even less meaningful<p>This is your opinion. Lots of tech money appears to disagree.<p>You should at least try to explain your contrarian position, or cite your favorite source that makes an argument that you believe.
                • regularfry2 hours ago
                  Lots of tech money is demonstrably wrong, at least in the general case. Microsoft&#x27;s whole business model is to be the second mover. Patents exist because first mover advantage alone isn&#x27;t enough to hold onto a win.<p>If AI is different, the onus is on AI folk to explain why, not the reverse.<p>The frontier labs are racing as hard as they can right now because they know as well as anyone else that they&#x27;re <i>at best</i> 6 months, probably closer to 3 months, ahead of the generic commodity AI market. The majority of the tech money currently boosting them is a bet that either they can stay that far ahead without tripping, or that a moat can be engineered. It&#x27;s not at all clear that it can - or, if it can, that it will before the money runs out.
                • ux2664788 hours ago
                  &gt; You should at least try to explain your contrarian position<p>FWIW I explained it fully and provided the argument, along with explicit grounding examples. Can you tell me what exactly you don&#x27;t understand?<p>To quote your other post:<p>&gt; This is the weakest part of the argument. Absolutely no reason to believe the second inventor will be 10x faster.<p>Nor is there any reason to believe that it will work at any usable speed. An LLM capable of AGI running at 0.000001tok&#x2F;s isn&#x27;t a very good head start. Nor is it a good moat if it can&#x27;t actually realize anything material fast enough. Lord knows America has labor problems.<p>I&#x27;m demonstrating the context is a lot more complicated than &quot;be first&quot;. There are innumerable specific examples. They&#x27;re not guaranteed to happen, that&#x27;s not the point being made.<p>Also &quot;no reason&quot; is curious when it&#x27;s the explicit pattern this exact industry has exhibited. Chinese models lag a few months, and come in swinging with an order of magnitude more efficiency. Does it apply? Who can say. It&#x27;s certainly on the table. Pretending like it&#x27;s any more ridiculous than AGI itself is irrational.<p>&gt; Does this ever happen? Even in traditional manufacturing, the second “inventor” starts behind and has to improve their own process to surpass the first.<p>To use the industry itself once again: Japan was the first to deep learning. Performance issues prevented them from capitalizing on it. Now they&#x27;re barely even a player on the field. Interestingly and adjacent to the industry, Japan was second to symbolic AI, and the FGCS was lightyears ahead of America&#x27;s crufty Lisp ecosystem. For a more modern example, Google was the first to transformers. Much of this revolution is thanks to them. A shame that Gemini is third-rate at best.<p>It&#x27;s not unique to this industry. That second inventor more often than not does a better job than the original one. It&#x27;s a pretty common pattern. Otherwise, we wouldn&#x27;t need patents.<p>My &quot;contrarian position&quot; is anything but. There is nothing new under the sun. You don&#x27;t just need to be first, you need to maintain exclusivity.
              • xboxnolifes10 hours ago
                Nobody argued exclusivity. But there have been clear advantages to having nukes first.
                • ux2664789 hours ago
                  &gt; Nobody argued exclusivity.<p>And I quote:<p>&gt; Imagine a world where only one country has AGI&#x2F;ASI.
          • xeromal12 hours ago
            Russia invaded Ukraine specifically because they have nukes. Everyone is afraid to fight back.
          • jr359214 hours ago
            As if time isn&#x27;t money? Everything is temporary... Getting somewhere first has immense value.
            • ux26647814 hours ago
              A reductive equation, economics isn&#x27;t thermodynamics. Money is fictional and value is subjective and unstable. Within this context, being first to AGI means nothing if the second invention of it comes 2 months later and works an order of magnitude faster than what the first iteration had self-improved to at that point in time. First mover advantage isn&#x27;t decisive, you have to actually be able to capitalize on it in a robust way.
              • fn-mote12 hours ago
                &gt; works an order of magnitude faster<p>This is the weakest part of the argument. Absolutely no reason to believe the second inventor will be 10x faster.<p>Does this ever happen? Even in traditional manufacturing, the second “inventor” starts behind and has to improve their own process to surpass the first.
              • 799tppp11 hours ago
                [dead]
        • expedited1232 hours ago
          &gt; Imagine a world where only one country has AGI&#x2F;ASI.<p>You&#x27;ve fallen for the hype. The only way these AI companies can justify their bullshit worth and burn of capital is by promising a magical AGI. It&#x27;s just marketing and it is still unclear if LLM&#x27;s will ever be profitable.<p>If they won&#x27;t - then EU is doing the right thing laying low.<p>&gt; Or a world where Europe only gets access to frontier models 6 months later.<p>What a calamity!
          • w4yai1 hour ago
            &gt; If they won&#x27;t - then EU is doing the right thing laying low.<p>Exactly. People seem to forget that US is burning money at an accelerating rate without any promise of returns. This is a high risk situation. Time will tell.
        • vrganj14 hours ago
          This presupposes AGI or ASI are a real, reachable thing. My read is they might be, but not as LLMs. Until there&#x27;s a fundamental rearchitecture, I&#x27;m AGI-agnostic and given that view, it doesn&#x27;t seem rational to bet the house on it.<p>We&#x27;ll see.
      • Palpatineli15 hours ago
        How about &quot;being able to align ASI somewhat to your values&quot;?
        • ben_w15 hours ago
          True, but that isn&#x27;t EU&#x2F;US&#x2F;China, it&#x27;s OpenAI&#x2F;Anthropic&#x2F;Grok&#x2F; …&#x2F;DeepMind (based in UK)&#x2F;… DeepSeek<p>With a lot of Chinese nationals in American companies, and an American corporation owning DeepMind, and a lot of people very upset with all of them at the same time, this is very messy.
      • tedggh14 hours ago
        A race to the bottom.
      • procgen15 hours ago
        The only prize is control of the light cone.
      • bpodgursky15 hours ago
        It&#x27;s red queen. You stay alive by winning, you lose everything by losing.
        • ben_w15 hours ago
          Not sure you get either outcome in either case.<p>Race dynamics increases p(doom) for everyone.<p>The non-doom scenarios include &quot;utopia for all&quot;, and &quot;power flows to investors, not citizens of whichever nation the winning model&#x27;s corp. was registered in&quot;.<p>Independently, &quot;oh look all the investors went bankrupt&quot; can happen in both &quot;doom&quot; and &quot;normal technology&quot; timelines.
      • ChrisClark15 hours ago
        Unless you manage to build a god, and keep it under control... okay, we&#x27;re all going to lose
    • podgorniy1 hour ago
      &gt; hope for Europe playing any significant role in the AI race<p>Ha. It&#x27;s a large market of the LLMs consumption. So it which will affect the AI race. Just from other perspective than you assumed.
    • barrenko2 hours ago
      They&#x27;ve been unfortunately hugged by the &quot;little death&quot; of working with the EU institutions, though I wouldn&#x27;t bet on them not coming back.
    • kolinko11 hours ago
      And US lost the significant role in chip&#x2F;pc manufacturing. But if a product becomes commodity or utility (which at least for now it seems is the direction), with little lockin, it&#x27;s not a big deal.<p>I hope we (EU) don&#x27;t waste money trying to train local models (which at least some people in Poland try to do), and tries to build our own chips - AI chips have different architecture than regular processor&#x2F;GPU, and TSMC doesn&#x27;t need to be winner in this new race.<p>And if not this, then smaller labs, harnesses and actual application.
    • pbkompasz14 hours ago
      Not being a rat in the rat race is the real win.
    • missedthecue12 hours ago
      The largest non-US non-China model is Russian which surprised me.
      • NicuCalcea12 hours ago
        Largest by what metric? And which model is that?
        • missedthecue8 hours ago
          Largest by total parameters. GigaChat 3.1 Ultra: 702B
      • sgt12 hours ago
        Are you saying Russia is winning over Europe in AI already? Dang, sanctions really don&#x27;t work.
        • IlikeMadison6 hours ago
          GigaChat is russo-centric and is completely useless outside Russia itself, performs absolutely worse than Mistral once you start using English. Reminder that Russia&#x27;s economy is smaller than Italy&#x27;s alone. They are not winning over Europe in any single metric, let alone against the poorest country in Europe, which is not even part of NATO nor the EU, on the battlefield.
          • sgt4 hours ago
            I&#x27;ve heard that before though. I don&#x27;t think you can compare economies like that since Russia has &quot;near infinite potential&quot; in natural resources, people, etc which Italy doesn&#x27;t.<p>They can kick into a true war economy and then become significantly more of a threat. They haven&#x27;t done so yet because the middle-class population likes their luxury lifestyle (for now).<p>I highly doubt Ukraine will &quot;win&quot; - in fact I believe Russia is ramping up to find an excuse at some point to attack a NATO country, but in a way and during a point in time it&#x27;s difficult for NATO to retaliate.
    • thadt15 hours ago
      Yeah? And here I&#x27;ve been a happy Transkribus customer for some time now. If there are better models or interfaces out there for analyzing historical handwriting, I&#x27;ll definitely take a look.
    • deadbabe12 hours ago
      Even Africa will surpass Europe with its massive data center upstarts breaking ground.
    • t3hTao13 hours ago
      [dead]
    • onetwig12 hours ago
      Huh? Mistral 7b was pioneering in its day and IMO they have been very on top of releasing niche useful models like moderation, OCR, etc.<p>I’m glad Mistral is working on useful solutions.<p>OpenAI&#x2F;Anthropic is like a retarded little sibling chasing “AGI” and giving up on rich media and other modalities.<p>OpenAI&#x2F;Anthropic is the worst of the mainstream AI.<p>It goes:<p>1. Gemini<p>2. Vidu<p>3. Le Chat (Mistral)<p>4. DeepAI<p>5. [insert MiniMax provider]
      • fn-mote12 hours ago
        Totally fake post. Gemini in the top?? For what?
        • v3ss0n9 hours ago
          Gemini is best at OCR and document extraction so far. We had tested for insurance needs. None match Gemini even flash level yet.
          • zzleeper7 hours ago
            Same here. Maybe Fable is better but in terms of cost effectiveness it wouldn&#x27;t even make sense to test it
  • fumeux_fume5 hours ago
    I think people misunderstand the utility of Mistral&#x27;s OCR. It&#x27;s not going to beat SOTA models for extraction on edge-case docs, but it&#x27;s MUCH cheaper and faster and does an excellent job on simple ones. I&#x27;ve been working on converting PDFs to EPUBs and Mistral has been making steady improvements. On a chapter of Bleak House it was able to extract and tag the header, titles, and references at the bottom every time. The only thing it struggled on was line numbers in the right margin which it correctly tagged as &quot;aside text&quot; 3&#x2F;5 times, but always separated from the core text each time. The important thing to keep in mind is that there&#x27;s no prompting needed, just upload the PDF and voila!. There&#x27;s even a batch mode with a 50% discount.
    • mkbkn1 hour ago
      &gt; The important thing to keep in mind is that there&#x27;s no prompting needed, just upload the PDF and voila!<p>I&#x27;m sorry, noob here. I have a special book that I bought which I can open only inside the Kindle app (Windows&#x2F;mobile). I have been meaning to screenshot the pages and convert them into a document&#x2F;PDF. What do I have to do to make it fast? Just upload all the screenshots one by one and tell Mistral &quot;Chat&quot; to OCR them?
  • waldrews12 hours ago
    The VLM&#x27;s are so good at complex document understanding now. But you just can&#x27;t trust them not to invisibly censor sensitive clinical&#x2F;legal docs, even at the maximally permissive settings.<p>And the deep learning OCR-only models won&#x27;t censor, but can and do hallucinate. I&#x27;ve yet to see a &#x27;scan with different approaches and reconcile and say you&#x27;re not sure if they don&#x27;t agree&#x27; system just work for generic complex documents.
    • kolinko11 hours ago
      The way my harness set it up is going through 2 or 3 providers, and cross-checking through them, also with plain text extracted if available.<p>I think we also had a layer that for any quote extracted tested it back if it exists within the original.<p>If you wanted 100% accuracy, I think it wouldn&#x27;t be too difficult nowadays to guess the font&amp;size&amp;other text settings, and re render the crucial parts.
    • dangoodmanUT10 hours ago
      &gt; But you just can&#x27;t trust them not to invisibly censor sensitive clinical&#x2F;legal docs<p>What&#x27;s an example of this?
  • merb16 hours ago
    1000 Pages &#x2F; 3.5€ this is expensive as hell. If this is not fastly superior than something like tesseract it is not worth it.
    • beernet15 hours ago
      Agreed. Does the GTM team there really sit together like &quot;oh yeah, that sounds reasonable&quot; while being totally beyond typical market prices?
    • Oras13 hours ago
      Even comparing to AWS Textract or Azure Document Intelligence, this is very expensive (more than double)
      • ttyyzz24 minutes ago
        How does it compare to those services? Are there any benchmarks yet?
  • piterrro15 hours ago
    For anyone interested, I have an ocr pipeline running on rented GPUs, doing around 1000pages for 0.05-01 usd with around 0.8 seconds per page with full bounding boxes support for grounding.<p>If you’re interested you can find contact to me via this profile.<p>3.5 usd&#x2F;1000 pages is just too expensive…
    • aliljet15 hours ago
      Accuracy is truly what people die for in the OCR game. Price isn&#x27;t the primary function here.. it&#x27;s an equation of price, accuracy, speed, and in mayn cases regulation.
      • merb15 hours ago
        Tbf even with tesseract you already get shit ton of accuracy and you can probably do these 1000 pages for way less than 3.5€. For 3.5€ you can spin up a cloud instance with 8vCPU+32gb on gcloud for 11 hours (or 11 instances for an hour) which can do way more than 1000 pages per hour on tesseract. It takes you around 6 second per page +-4 seconds start&#x2F;stop depending on what you are doing on that instance size without too much optimization (you can probably even run multiple processes on a single node)<p>Google documentai costs 1.5$ per 1000 which is probably better in quality and speed.
        • kergonath14 hours ago
          Tesseract is not a substitute for these models, which understand complex layouts and also extract bounding boxes for things like tables and pictures. They are also much better at making sense of cursive scripts.<p>I’ve been there, implementing a way to linearise text from a document with pages with 1, 2 or 3 columns, some of them in landscape is a nightmare. And that’s not even considering equations.<p>In the end it’s way easier to use a specialised model, trained by other people to do exactly what I need.
        • tjoff13 hours ago
          Tesseract is super picky though often failing pixel-perfect screenshots...<p>So, maybe it can be tuned for your usecase but with that kind of investment €3.5 for 1000 pages is a <i>bargain</i>...
          • merb12 hours ago
            It‘s also cheaper than the murican ones from Google, Amazon, …. And tesseract was an example. Heck you can go xberg and use paddleocr. Most often layout is less of a problem for ocr. Most often you need high accuracy, which tools like these are often worse in the 95 percentile.
        • v3ss0n8 hours ago
          Are you serious? Tasserect is the worst compare to PaddleOCR , Surya&#x2F;Marker
      • piterrro14 hours ago
        Most use cases dont need that kind of accuracy, just doesnt justify the 3-4usd range. I build for that exact case (tender documents, we’re processing north of 100k pages per day), it doesnt need to recognize scanned written text from 1930s, its usually pdf&#x2F;docs&#x2F;scanned printed pages. The accuracy is great, bounding boxes are must have for proper grounding for building answers by LLMs. Tesseract was too slow and not enough in some cases (for example tables or images which we also recognize and describe)
        • pogue14 hours ago
          If you&#x27;re getting inaccurate results from OCR what&#x27;s the purpose of even doing it? Inaccuracy of text of any kind seems like a completely obvious failure of the entire purpose of scanning text into a computer.
          • fluoridation14 hours ago
            It depends on what you need. For example a while ago I scanned and OCR&#x27;ed a bunch of receipts to get a timeline of my salary. I only cared about the gross and net figures, and nothing else mattered. Tesseract&#x27;s output had a bunch of errors and misdetections, but the main figures always came out OK, and a local LLM was able to pick them out from the noise every time.<p>There&#x27;s a big gulf between &quot;it&#x27;s as if a human being had transcribed it and reconstructed the original document&quot; and &quot;so completely broken it can&#x27;t be used for anything&quot;.
          • piterrro14 hours ago
            Accuracy can have different dimensions, depends on what you can tolerate and whether you can detect it to apply more powerful methods.<p>Imagine you have a cheap and 99% accurate ocr. The other 1% you can detect and apply more powerful (more accurate but slower and more expensive) ocr method. What would you use? At scale these things add up.
    • Telemakhos13 hours ago
      Can it produce accessible PDF files that will pass accessibility tests? Someone who can do that will make a killing laundering PDFs for academia: by April 26, every PDF, syllabus, and academic document needs to comply with WCAG 2.1 Level AA, which means structural tagging, alt text, and lots of other checklist items that AI could probably generate.
    • x3ro15 hours ago
      You should put contact details in your profile :)
    • vrganj14 hours ago
      Is it European-hosted and fully outside of both CLOUD Act and CCP reach?<p>Because I&#x27;m assuming that&#x27;s why they get to charge more for the right type of customer.
      • piterrro14 hours ago
        You can even run it on your desk if you want, a single gtx 4090 is enough. It can be fully air gapped.
  • einpoklum11 minutes ago
    How does this perform with non-Latin and non-LTR scripts? Say, Chinese, Arabic, Devangari, Adlam, etc.?
  • ks204813 hours ago
    Does anyone know a site that lets you browse examples of input &#x2F; output pairs?, particularly with layout analysis (bounding boxes of figures, tables, etc).
  • ianhawes15 hours ago
    I won&#x27;t comment on accuracy, but in internal benchmarks, Mistral OCR is significantly faster than comparable APIs.
  • Johnny_Bonk15 hours ago
    How does this compare to Baidu Unlimited OCR. I&#x27;ve been very impressed with Baidu and it&#x27;s essentially free to run on a decent computer, other than electricity costs.
    • spiderfarmer15 hours ago
      Where do your documents go?
      • rescbr15 hours ago
        They go to the decent computer hosting the model, which can be yours if you pay the electricity costs
  • ad_fontes15 hours ago
    I&#x27;ve been experimenting with using NuExtract this week on locally OCRing bank statements that don&#x27;t have a predefined document structure. It&#x27;s <i>way</i> better than Tesseract or a generic vision-enabled model. It runs great on a single RTX 4090 at the modest throughput I need.<p>Their hosted, API-based service is something like a third of the cost of this model.
  • maelito13 hours ago
    Given the latest vibe release&#x27;s new &quot;follow default&quot; model option, we should see a new coding &#x2F; general Mistral model, mistral 4, soon.
  • oliveralbertini12 hours ago
    I&#x27;m wondering if this model performs better on french (and other european languages) documents than others
  • parhamn14 hours ago
    Mistral is bumping the price of this thing every release. I think we&#x27;re at 2x now?
  • maz1b15 hours ago
    How does this compare to 4?
  • mainecoder16 hours ago
    The chinese did it better, mistral is alive thanks to regulations.
    • eloisant11 minutes ago
      It&#x27;s not regulations, it&#x27;s European companies being (legitimately) worried about their data sovereignty if they use US or Chinese AI providers.
    • maelito14 hours ago
      No it&#x27;s the other way round : the Chinese do better thanks to regulations : massive amounts of money from Big tech and public money.
    • Bombthecat14 hours ago
      Yeah, I&#x27;m not sending personal bills etc to china. No thanks
      • gkbrk13 hours ago
        It&#x27;s still open-weight models that you can download and run locally. You don&#x27;t need to send anything to China.
      • BlackRabbit14 hours ago
        You can use from EU&#x2F;US hosters. No need to send data to China.
      • sajithdilshan12 hours ago
        and you trust the french?
        • MrDrMcCoy9 hours ago
          I trust them more than the Americans and Chinese. Genuinely rooting for them to close the gap.
          • sajithdilshan35 minutes ago
            Good luck with that. Mistral is never going to produce a frontier model. Their business model is to rely on the EU regulations and cater to the bureaucracy. I won&#x27;t be surprised if their next model is specialized on generating laws&#x2F;regulations for EU lawmakers so they can sell it to EU. That way their business would survive for the next 5-10 years till EU implodes.
    • mangecoeur15 hours ago
      i.e. it&#x27;s one AI company that&#x27;s basically guaranteed to never fail since it has a market niche guaranteed by European companies and governments.
      • petcat15 hours ago
        Which is also why their most recent model &quot;Shieldstral&quot; does nothing except monitor and moderate internet content.<p>After stuff like Chat Control I think they&#x27;re obviously seeing a big demand for this kind of &quot;internet safety&quot; technology in Europe.
    • rtaylorgarlock15 hours ago
      I&#x27;ve been a bit more careful about complaining about regulations broadly due to competitive advantage, e.g. ITAR
    • t3hTao13 hours ago
      [dead]
  • tethys12 hours ago
    Who the hell at Mistral thinks it is a good idea to register CMD + T as a shortcut for switching theme!?
  • nc55g3g6 hours ago
    [flagged]
  • t3hTao13 hours ago
    [dead]
  • hmokiguess14 hours ago
    Whoever is paying all that for OCR is being scammed.