Kimi K3-256k

(kimi.com)

349 points by monneyboi7 hours ago

24 comments

  • xyzsparetimexyz5 hours ago
    Wow. So kimi is suddenly half the price for all users until they hit 256k of context? Thats massive.
    • InsideOutSanta5 hours ago
      I don&#x27;t think so. This is a separate model, so I assume that if you just use this and switch to the 1 million context model when you reach 256k, your cache will be invalidated, so you&#x27;ll re-pay the 256k tokens on the 1 million context model pricing.<p>Edit: I was wrong, thanks to longwave for pointing this out. It&#x27;s absolutely possible to start out on the 256k model and then switch to the 1 million model when you get close to the context limit without invalidating the cache:<p><i>&quot;When switching from k3-256k to k3 (1M), if k3-256k is close to the 256k limit and you don&#x27;t want compact to lose information, you can switch directly to 1M. The current version switching from 256k to 1M does not affect the cache.&quot;</i>
      • longwave4 hours ago
        The article explicitly says &quot;The current version switching from 256k to 1M does not affect the cache.&quot;
    • conradludgate4 hours ago
      As far as I understand, no. They&#x27;re suggesting that smaller context windows are typically cheaper (fewer input tokens over time).
    • StopTencent1 hour ago
      [flagged]
  • wren69915 hours ago
    This seems functionally similar to OpenAI having a step in pricing once you exceed a certain context length (also at 272k aka 2^18 aka 256k).<p>Having a lot of active context increases the per-token cost (flops issued and bytes read per token out) so it makes sense to pass that cost on to users. I&#x27;m actually surprised it&#x27;s implemented as a hard cutoff instead of a smooth gradient.
    • pornel50 minutes ago
      RAM needed for keeping KV cache around may be the more expensive factor.
    • BoorishBears5 hours ago
      Not surprising it&#x27;s a hard cutoff: they almost certainly have two infrastructure configurations for the two max sequence lengths<p>Fewer nodes dedicated to prefill per instance, and fewer nodes in total since you don&#x27;t need to support a higher KV cache.<p>Disaggregated inference also means they can tune the balance of compute dedicated to prefill seperately from decode
      • MrBuddyCasino1 hour ago
        AI infra buildup is so massive that the frontier labs should be able to offer more than one level of context length to incentivize token thriftiness.<p>One would think compute-constrained actors like Anthropic would have done so, unless prefill isn’t really a bottleneck compared to decode?
  • hawtads6 hours ago
    This is just an API level change right? The model itself should be the same I think.
  • wxw6 hours ago
    &gt; k3-256k is now available. Within 256k context, it delivers the same results. k3 (1M) consumes about twice as much quota as k3-256k.
  • dgritsko6 hours ago
    This isn&#x27;t quantized, right? Just a smaller context?
    • DSingularity6 hours ago
      Its 256k context window. Quantization is orthogonal. We cant really tell directly so it could be quantized.
  • try-working5 hours ago
    i&#x27;ve never had any issues with 256k context. see no reason to bump up to 1m if it comes at a premium.
    • trollbridge1 hour ago
      It&#x27;s pretty handy for have very long contexts for long running agents, or else when doing literary analysis to simply be able to load the entire book in.
  • dools5 hours ago
    Hopefully this helps reduce some of the pressure on their infrastructure. Their models have all become super dumb recently and their support are not addressing it. I have a hunch they’ve been serving a significant percentage of requests with quantised models.
    • stingraycharles4 hours ago
      Ah come on, HN is not the place for these types of unfounded conspiracy theories that keep popping up on Reddit.
  • lukan6 hours ago
    Since Claude is the first time for me really, really out (TIL against my wished about <a href="https:&#x2F;&#x2F;status.claude.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;status.claude.com&#x2F;</a>), I am now interested enough to see what else works. But ... when I click pricing, I see &quot;Join a waitlist&quot;. Wtf? Are they really that good, so were totally surprised and overwhelmed by the requests, is this a marketing stunt, or do they just don&#x27;t have the hardware being in china?
    • KronisLV6 hours ago
      &gt; Are they really that good,<p>They did exceed the expecations of pretty much everyone! I&#x27;ve also blogged about using the model, it&#x27;s a bit on the slow side but pretty good!<p>&gt; so were totally surprised and overwhelmed by the requests ... or do they just don&#x27;t have the hardware being in china?<p>Yes, this is mostly the case: <a href="https:&#x2F;&#x2F;x.com&#x2F;Kimi_Moonshot&#x2F;status&#x2F;2078855608565207130" rel="nofollow">https:&#x2F;&#x2F;x.com&#x2F;Kimi_Moonshot&#x2F;status&#x2F;2078855608565207130</a><p>As a user, I much prefer that to service disruptions or severely degraded or secretly quantized performance. However if I didn&#x27;t have an account, I&#x27;d be pretty pissed off about not being able to give them money and become a user.
      • lukan5 hours ago
        &quot;I&#x27;d be pretty pissed off about not being able to give them money and become a user.&quot;<p>I am now rather pretty pissed towards antrophic for stopping my flow and forcing me to search for alternatives.
    • MYEUHD6 hours ago
      Since the model is open-weights, you can get from other providers, for example see <a href="https:&#x2F;&#x2F;openrouter.ai&#x2F;moonshotai&#x2F;kimi-k3#providers" rel="nofollow">https:&#x2F;&#x2F;openrouter.ai&#x2F;moonshotai&#x2F;kimi-k3#providers</a>
      • jadbox15 minutes ago
        Do any support the smaller context for better pricing?
      • Alifatisk6 hours ago
        The only downside with third party providers is that you have to trust that the provider have setup and configured it correctly, and is not secretly quantizing it.<p>See Kimi Vendor Verifier.
        • usef-4 hours ago
          Plenty of providers on that list aren&#x27;t fly-by-nights and have easily proven themselves in past models. Jeremy Howard does demos on Fireworks.<p>A nice thing about open router is you can specify filters, like &quot;US hosting without data retention&quot;
          • applicative49 minutes ago
            Yes, something I read just before it was released suggested that, because of the unique features of the model, there was a back and forth of Hugging Face, Moonshot and providers like Together AI and Fireworks AI. This also explained why it took much less than a day for Together, Fireworks etc to appear. Whatever they are doing is what Moonshot wants, I think.
    • HDBaseT4 hours ago
      Not at all a marketing gimmick, the demand is simply that high.<p>I was able to press &quot;Join waitlist&quot; and then within 48 hours got accepted. The limits aren&#x27;t very high, no where near the endless subsided+resets given on ChatGPT&#x2F;Claude. I recommend ChatGPT for good value output!<p>Others here mentioned &quot;the providers could be quantizing it!&quot; but some of the providers on OpenRouter have partnered with Moonshoot and OpenRouter shows the int when you expand on the provider.<p>It should say &quot;mxfp4&quot; but providers like Baseten report FP8.
    • Alifatisk6 hours ago
      No, the waiting list is true.<p>Kimi had become that popular. I was a subscriber of Kimi back when latest version was Kimi K2. Later I unsubscribed because I jumped over to GLM subscription (they had amazing deal). Now when I wanted to try out Kimi K3 to find out what the fuzz was all about, I couldn’t subscribe to them.<p>I remember reading a post from Moonshot team about this, they are doing this because they are almost at peak capacity and want to reserve it to keep the quality for their current customers.<p>We are actually witnessing an open-weight model catching up at catching mainstream users attention. And instead of behaving like Anthropic, they actually care about their users experience.
      • therein5 hours ago
        I wish I could get a Kimi subscription. I&#x27;d jump away from Anthropic in a heartbeat.<p>When they first announced, I created an account but didn&#x27;t subscribe, so I am stuck on a waitlist now.
    • InsideOutSanta6 hours ago
      It&#x27;s an extremely popular model hosted by a company affected by hardware export bans. I doubt they&#x27;d voluntarily prevent people from subscribing if they didn&#x27;t absolutely have to to maintain service quality.
      • zorked5 hours ago
        I have been using Kimi for months. It did get superslow after the K3 launch, then they blocked signups and normalized the service.<p>It&#x27;s not ideal that people can&#x27;t join but as a user I am happy that they focused on serving existing customers.
  • timcobb4 hours ago
    Codex uses 256k masterfully, 1M is luxurious but still quite expensive and seems not necessary as a default.
    • erichocean5 minutes ago
      That&#x27;s because it has the best compaction of anyone, and it&#x27;s not even close.<p>Claude might as well not even do it in my experience.
    • cybersec4732 minutes ago
      &gt; Codex uses 256k masterfully<p>This.<p>When working on hard problems (not &quot;vibecode me a script to show an alert box&quot;, but e.g. &quot;let&#x27;s see what this three-level LUT-state-machine obfuscated binary does&quot;), hitting the 1M (!) context window with Claude Code feels like you were talking to Claude Claudewski when his shift just abruptly ends, he packs his things, throws the office keys at Claude Claudeson in-between the front door frame while handovering like &quot;Hi! Nice to see you, good luck.&quot; and now here we go again, you are working with someone who just experienced an acute amnesia. It tries everything it already tried, everything it was told in the initial prompt to not do, everything it was told in follow-up prompts not to do. &quot;You were right, this approach does not work and we don&#x27;t have 20 TB RAM on this machine for full symbolic execution, let me try...&quot;<p>In Codex, it&#x27;s so seamless that I sometimes just notice &quot;wait, the context was 20 % remaining, it is 70 % now, wow, when did this happen&quot;, while it seamlessly works on the task. Basically never had an issue with context on Codex, be it coding features, cracking hard crack-me ciphers, or researching basically anything.<p>(For full disclosure, my last experience with Claude was a few weeks ago when I cancelled the subscription, maybe they fully reworked the traumatic &quot;Summarizing&quot; - &quot;Oh, hi! Where are we? Who am I? What we are doing? This is taking too long, let me take a shortcut...&quot; lobotomy they were doing in the meantime.)<p>256k is enough when the harness uses it properly and the model is not stupid. And also when the tokenizer is not tuned to invoice as many tokens as possible...
    • jsemrau2 hours ago
      Needing large context windows is an illusion.
    • Tostino1 hour ago
      I&#x27;ll have to give it a shot some time soon. I just can&#x27;t imagine working on any of my serious projects doing a new large feature with that little room. By the time it has looked at half the code required to start planning the work, it&#x27;d be out of context. Must use sub agents better or something.
      • cybersec4721 minutes ago
        Also, a token means something different in each model&#x2F;tokenizer. &quot;For improved performance we will use (read: invoice you) 30 % more&quot;...<p><a href="https:&#x2F;&#x2F;platform.claude.com&#x2F;docs&#x2F;en&#x2F;about-claude&#x2F;models&#x2F;whats-new-sonnet-5#new-tokenizer" rel="nofollow">https:&#x2F;&#x2F;platform.claude.com&#x2F;docs&#x2F;en&#x2F;about-claude&#x2F;models&#x2F;what...</a>
  • madihaa7 hours ago
    That&#x27;s actually nice! I usually try to stay below 200k context anyway.
    • chipgap985 hours ago
      This exact same comment is the top comment on the reddit thread for this news<p><a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;kimi&#x2F;s&#x2F;BFa1TR9vNg" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;kimi&#x2F;s&#x2F;BFa1TR9vNg</a>
      • Permit4 hours ago
        I searched one other comment of theirs in their history and see the same thing. On mobile so I can’t easily tell if this account is the bot account or if the ones on Reddit are.<p>HN: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47934437">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=47934437</a><p>Reddit: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;worldnews&#x2F;comments&#x2F;1sxzzop&#x2F;comment&#x2F;oiqkl33&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;worldnews&#x2F;comments&#x2F;1sxzzop&#x2F;comment&#x2F;...</a>
        • throw1092012 minutes ago
          The Reddit comments are both older.
        • altern82 hours ago
          Karma-farming bot..?
      • Barbing4 hours ago
        Does that URL link back to you somehow?<p>Full: <a href="https:&#x2F;&#x2F;reddit.com&#x2F;r&#x2F;kimi&#x2F;comments&#x2F;1v9aqsi&#x2F;comment&#x2F;p0cbe0n&#x2F;" rel="nofollow">https:&#x2F;&#x2F;reddit.com&#x2F;r&#x2F;kimi&#x2F;comments&#x2F;1v9aqsi&#x2F;comment&#x2F;p0cbe0n&#x2F;</a><p>Mirror: <a href="https:&#x2F;&#x2F;redlib.us.catsarch.com&#x2F;r&#x2F;kimi&#x2F;comments&#x2F;1v9aqsi&#x2F;k3256k_is_now_available&#x2F;p0cbe0n" rel="nofollow">https:&#x2F;&#x2F;redlib.us.catsarch.com&#x2F;r&#x2F;kimi&#x2F;comments&#x2F;1v9aqsi&#x2F;k3256...</a>
    • giancarlostoro6 hours ago
      For me the sweet spot is somewhere under 500k depending on how extensive I want to get. You can build up a sizable effort project in half a million tokens with Claude, with Claude having all the context from ground 0 to wherever you&#x27;re off at.
      • cyanydeez6 hours ago
        I&#x27;m always curious what you guys are working on; every git repo I&#x27;ve run a local model on and stick below &lt;100k to increase speed seems effective enough to scope patches and changes.
        • KronisLV6 hours ago
          My current Claude Code session has been going on for like 35 hours and has used up around 400 million tokens, thankfully almost all of those being cached (95-98%) - pretty typical for long form agentic work.<p>First you spend like 2-3 hours working on a plan, once you have that you just tell the model to go and implement it, do adversarial sub-agent review loops before each commit and also make sure that all tooling and tests pass (including coverage requirements). You do need to poke it in a slightly different direction every few hours, though. Not even any novel work, just some refactoring and SSE notification hardening, bug fixes, alongside environment tuning and getting rid of some bottlenecks (also migrated from Oracle to PostgreSQL but that&#x27;s mostly done).<p>That said, Kimi somehow manages to use less context in the main thread than Anthropic&#x27;s models (even when you use sub-agents and also dynamic workflows in Claude Code), might have something to do with either how the model is tuned or their Kimi Code harness - because even in most of the longer form sessions it doesn&#x27;t seem to fill up quite as quickly (note: because the kimi vis tool doesn&#x27;t have a full summary view across all agents, these are the main long running agent stats across some sessions, not sub-agents):<p><pre><code> total tokens cache hit rate wall time peak context 283M 98% 3963m 466k 258M 97% 2724m 467k 98M 94% 1353m 393k 67M 97% 614m 434k 75M 98% 1447m 498k 53M 99% 191m 375k 6M 96% 139m 124k 7M 98% 86m 118k 11M 99% 61m 147k </code></pre> I could see 256k context being sufficient for all sorts of work, even if intermediate progress&#x2F;plan tracking files and docs might have to be used along the way, in addition to whatever plan support the harness has (for example, if you document something that will be relevant for load testing you might need that in 10 turns but not during the ones before then).
          • smotched5 hours ago
            Once you have the plan you don&#x27;t need to keep the 2hrs of research in the context (which is most of it) you can drop that plan into a file and start fresh for implementation.
            • XenophileJKO4 hours ago
              That is true, but I feel like sometimes if the conversation contains useful rational it can help to keep it. I think sometimes it is a judgement call, I will sometimes compress the context first.<p>If I feel like the model and I explored a lot of options I won&#x27;t want to keep context as it might be confusing.<p>I think the more you use it the better judge you are of whether you should purge, compress, or just keep the context before executing the plan.
              • dandaka4 hours ago
                Models work best when they have short instructions and no noise. &quot;Conversation [may] contain&quot; also means &quot;conversation has a lot of noise&quot;. It degrades performance and increases cost.<p>I use handoff skill to ask model to write a prompt for itself.
                • pertymcpert4 hours ago
                  The applicability of your advice is very model dependent. Some like Claude have very good long context performance, whereas others they fall off much quicker past some threshold.
          • bkaae6 hours ago
            Thanks for sharing - is this a normal feature request you are implementing in this example or is this a project from scratch? Trying to get an idea of how your workflow compares to mine.
            • KronisLV3 hours ago
              Existing project and as usual, a few issues mashed together in one mostly coherent plan. The shorter Kimi sessions I mentioned were singular features, there 256k would be wholly adequate.<p>A greenfield project would probably allow at least 2x fewer tokens to be used in most of those long tasks, but I was mostly after consistency and bug fixes along the way as needed.
          • vidarh4 hours ago
            Kimi CLI has a mechanism with checkpoints and the ability for the agent to revert to a checkpoint + a message of how to continue based on what went right&#x2F;wrong. I don&#x27;t know if that&#x27;s the cause of what you&#x27;ve seen, but it&#x27;s plausible.
            • DelightOne4 hours ago
              With prefix caching, you get checkpoints for free. Do you mean that?
              • vidarh3 hours ago
                No. Prefix caching is just an optimisation on the server side. What Kimi CLI does is insert &lt;system&gt; tags that include a checkpoint marker with an id.<p>The model is then given a tool that allows the <i>model</i> to decide to roll back to a checkpoint + a message containing any additional useful information.<p>It&#x27;s specifically instructed to use that tool[1] in cases like when it has inadvertendly read a large file where most of the content is not relevant to the task, or after a web search where it&#x27;s found what it&#x27;s looking for but most of the content isn&#x27;t needed, or when it&#x27;s written code that didn&#x27;t work as expected, or similar.<p>It basically lets the model backtrack and &quot;forget&quot; irrelevant details at the end of the context but give itself hints on how it should continue from the checkpoint.<p>Though, interestingly they seem to be abandoning it in their new CLI (kimi-code), unless it&#x27;s been folded into other functionality. Not sure if they just feel it&#x27;s not needed any more with their newer models or if it just didn&#x27;t work as well as they expected.<p>[1] named &quot;D-Mail&quot;, or &quot;DeLorean Mail&quot; in a reference to Steins;Gate, which again references Back To The Future. See <a href="https:&#x2F;&#x2F;github.com&#x2F;MoonshotAI&#x2F;kimi-cli&#x2F;blob&#x2F;main&#x2F;src&#x2F;kimi_cli&#x2F;tools&#x2F;dmail&#x2F;dmail.md" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;MoonshotAI&#x2F;kimi-cli&#x2F;blob&#x2F;main&#x2F;src&#x2F;kimi_cl...</a> and <a href="https:&#x2F;&#x2F;steins-gate.fandom.com&#x2F;wiki&#x2F;D-Mail" rel="nofollow">https:&#x2F;&#x2F;steins-gate.fandom.com&#x2F;wiki&#x2F;D-Mail</a>
        • giancarlostoro6 hours ago
          I had Claude build me a Python-inspired .NET language that treats .NET as a first class citizen, and breaks backwards compatibility where some Python nuances don&#x27;t really apply to .NET for. I was able to get it to build a sample ASP .NET Web application that ran on Culebral code.<p>Haven&#x27;t gone back to it, have been using Claude Code on a private project I&#x27;m still architecting.<p><a href="https:&#x2F;&#x2F;github.com&#x2F;Giancarlos&#x2F;Culebral" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Giancarlos&#x2F;Culebral</a>
        • _zoltan_5 hours ago
          I have a couple projects where the background research is easily over 500k without writing any code, after ultracode subagents synthesis.
          • giancarlostoro5 hours ago
            I think the distinction is when the model decides to read code top to bottom vs when the model chooses to parse code indirectly to save on tokens, there&#x27;s also AST tooling to let the model see project structure.
        • sheeshkebab6 hours ago
          Try doing a refactoring of some sort or larger new feature using just an agent on a moderately sized codebase, 256k will be compacting every few minutes, and result will be unusable.
          • hedgehog5 hours ago
            I do essentially all of my work with auto compact set to 250k and it&#x27;s fine. It may be due to the way the project tooling is set up and the use of sub agents?
        • jdoe1337halo6 hours ago
          They are just talking to the model in CC, while staying in a single thread. Doubt they have any actual coding knowledge to compartmentalize different problems in the codebase.
          • giancarlostoro6 hours ago
            Depends on the programming language I&#x27;m using for a given project, and the domain I&#x27;m working with. I&#x27;ve been coding as a hobbyist for nearly two decades now (since my teens), professionally for 9 years, and was a TA before that for roughly 3 years at one of the best colleges for this field in the state (at least back then it was) where I taught other students about programming, in some cases I was their primary learning resource.<p>But yeah, I have no idea about anything about software because you made an assumption off very little to go by.
    • pesfandiar5 hours ago
      &quot;256k ought to be enough for anybody&quot;
  • MangoCoffee5 hours ago
    LLMs is quickly became commodities. US AI labs like OpenAI is losing their moat. Hyperscalers and data center owners who can sell cheap token will win
    • okrad5 hours ago
      I believe Codex harness is pretty sticky though I haven’t tried many others. Does anyone else provide a harness of that quality?
      • nvarsj4 hours ago
        Pi harness with subagents is pretty great. I’m using it for everything now.
        • yeswecatan2 hours ago
          How are you handling subagents?
          • anderber2 hours ago
            Usually the main agent can handle creating subagents
      • conradludgate4 hours ago
        Their harness is indeed nice, but we&#x27;re using it against our own AI Gateway at work. Right now we only expose GPT models on it, but I imagine it&#x27;s possible to add our own models eventually
      • verdverm1 hour ago
        opencode is highly regarded, I use it exclusively
    • ux2664784 hours ago
      I have so much gratitude to the frontier companies who did all the extremely complicated research and development, model by model. It already feels difficult to remember how much capital it really took. Thank you for getting us to this point.
  • gigatexal4 hours ago
    Not relevant to this link but I was thinking about the allegations of Chinese AI companies distilling from the big frontier American ones. And I came to the conclusion: I don’t care.<p>Who cares? China has always copied and then copied the means of production and then out produced. See also Tesla and now all the Chinese cars eating their lunch.<p>As long as I get really solid AI models for cheap that do what I need I don’t care if they’re Chinese or otherwise.<p>I’ll still never use Grok from SpaceX AI cuz eww no, I have principles. ;-)
  • jedisct15 hours ago
    This is a fantastic option for swival.dev given its very efficient context management compared to e.g. Claude Code.
  • jscott8175 hours ago
    Can we assume that model performance at 90% of the 256k limit != 90% of 1M token limit?<p>Is this the exact same model just with less VRAM allocated for context window?
  • hendersoon6 hours ago
    What is the purpose of this? Just a hard cutoff below the actual context window? You could set that in your harness anyway.
    • meatmanek4 hours ago
      At least in the self-hosted LLM inference engines, you have to pre-allocate space for the maximum amount of context you want to allow for each parallel session. By using a lower maximum, you don&#x27;t have to allocate as much VRAM for each session, allowing more usage for the same amount of hardware. Thus, cheaper.
    • dnlzro6 hours ago
      Uses less quota (i.e., cheaper). For people who like to keep their contexts small, this is a no-brainer.
    • surgical_fire6 hours ago
      &gt; k3 (1M) consumes about twice as much quota as k3-256k<p>Cheaper?
      • chrisweekly5 hours ago
        k3-256 consumes half the quota, so... yes? (assuming the user isn&#x27;t making use of the larger context window)
  • sergiotapia6 hours ago
    I can&#x27;t seem to find pricing for this model. Since the context size is just a quarter of the full size K3, is the price also much cheaper?<p>I usually keep my context in chats below 256k anyways so this would be tremendous honestly.
    • daemonologist6 hours ago
      It seems to only be available in Kimi Code, via subscription, no there&#x27;s no API pricing. The linked page says it consumes about half as much quota as the 1M version though.
  • Conol_ai4 minutes ago
    [flagged]
  • Conol_ai21 minutes ago
    [flagged]
  • 1saadcodes3 hours ago
    [dead]
  • illithid06 hours ago
    This was posted 38 minutes ago, and as of 20 minutes ago, several Anthropic services are now designated as having a &quot;major outage&quot;.<p>Doubt these are related, but it made me laugh a little.
    • paxys6 hours ago
      Anthropic services have outages on all days ending in y.
      • de6u99er5 hours ago
        So users in Germany are not affected?
        • jampa4 hours ago
          Anthropic has better SLA in Germany. I’ve heard uptime there can get up to nein nines.
        • zarmin5 hours ago
          Germany ends in y
          • tshaddox5 hours ago
            Not in Germany.
            • hebelehubele5 hours ago
              The word &quot;Germany&quot; is &quot;Germany&quot; everwhere and always ends with y. The country Germany has different names though.
              • aksappy5 hours ago
                Germans prefer Deutschland, no?
                • 3619947524 hours ago
                  so Deutschland is not affected. Germany is
          • vidarh3 hours ago
            But Tag doesn&#x27;t.
  • superloika6 hours ago
    [flagged]
  • ibuildproducts6 hours ago
    omg! new model!!
    • Alifatisk6 hours ago
      Same model, new configuration
  • holoduke6 hours ago
    A bit of topic. But how likely is it that the US will restrict Chinese open weight models and also force Euro countries to do the same? I think it will be effective within 6 months. The US is having a hard time staying competitive.
    • jeppebemad6 hours ago
      I don’t know, but I do think that the days of the US “forcing” Euro countries to do anything, is over.
      • holoduke5 hours ago
        No it&#x27;s not. The US dictates every single step in Europe. Euro politicians are in the pockets of American institutions. The biggest reason of dumb decisions in the EU is because of deliberate decisions made in favour of the US. LNG, war in Ukraine, ASML export restrictions and many more.
    • impossiblefork5 hours ago
      It isn&#x27;t legally possible for them to do this at the EU level. The EU parliament would never vote for it.<p>For pressure at the country level leading to this kind of thing I think it&#x27;s very unlikely. Here in Sweden it wouldn&#x27;t just require a vote in the Swedish parliament and before this there&#x27;d have to be förarbeten and you can&#x27;t just brazenly push things through with insane arguments, Swedish social convention goes against it-- and there&#x27;s just no way to get it through.<p>It also might not even be legal. &quot;We aren&#x27;t at war with China and I&#x27;m a communist, and the US LLMs are so aligned with values inimical to my political ideology that this is interference with opinion formation&quot; might be an actual legal argument that the ECHR or CJEU might actually have to accept.
      • holoduke5 hours ago
        It happened with ASML. What&#x27;s your thought on it? The US forbid dutch company to export.
        • impossiblefork4 hours ago
          Well, that&#x27;s the deal, I assume-- that they weren&#x27;t allowed to buy the Japanese light source outright, so they bought the American one, even though it required giving the Americans some sort of veto or control.<p>I guess it sucks if one wants to expert broadly, but if you&#x27;re big on vertical integration and the Japanese won&#x27;t sell I guess you take what you can get.
    • realusername5 hours ago
      Trump burned a lot of bridges in the EU, that one will be a hard sell
  • periodjet5 hours ago
    Why are Anthropic and OpenAI even allowing their coding harness apps to be plugged into different model providers…? I’m surprised they haven’t figured out a way to clamp down on that by now.
    • wren69915 hours ago
      I&#x27;d guess because it costs them nothing and it gives you a smoother transition back towards paying for their products.
    • archildress5 hours ago
      From my view, as soon as they do that, they send people out the door to use Opencode instead - and once many people have a taste of trying every model via Openrouter, it&#x27;s eye opening as to the possibilities.<p>Of course - Anthropic and OpenAI have an advantage in the amount they can subsidize the usage, but I think those days are waning.
    • auspiv5 hours ago
      Well when a huge part of potential revenue is all in on Bedrock... you need the harness to be able to talk to Bedrock. And Vertex. And all the other places these models are hosted. And allow for proxy because many businesses do not all direct internet access... all valid business reasons.
    • HDBaseT4 hours ago
      It honestly should be easier, its a by product of everyone using the same API standard.<p>Codex and Claude require editing a .json file, but most other harnesses have direct connections via a &#x2F;provider or &#x2F;login command.