Rio de Janeiro's city government model Rio3.5 beats Qwen3.7 in recent benchmarks

(twitter.com)

111 points by lucasfcosta3 hours ago

12 comments

VoidWhisperer1 hour ago
<a href="https://github.com/nex-agi/Nex-N2/issues/4" rel="nofollow">https://github.com/nex-agi/Nex-N2/issues/4</a>Seems that they didn't make/train a new novel model, they did a mix of two existing models and then gave it an instruction to say it was 'Rio, trained by Rio AI Labs'
- w4yai1 hour ago
 > The model is built via a merge of <a href="https://huggingface.co/nex-agi/Nex-N2-Pro" rel="nofollow">https://huggingface.co/nex-agi/Nex-N2-Pro</a> and <a href="https://huggingface.co/Qwen/Qwen3.5-397B-A17B" rel="nofollow">https://huggingface.co/Qwen/Qwen3.5-397B-A17B</a>, proceeded by On-Policy Distillation from a stronger model. We detected an incorrect upload in the previous version, where the base merged version was upload instead of the final distilled model. We are sorry for the confusion and apologize profusely.<a href="https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B/commit/a778c1ec4e21180ee55c3ea016a348e549e75f09" rel="nofollow">https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B/comm...</a>
 - daquisu8 minutes ago
 It was a recent edit though. Yesterday snapshot: <a href="https://web.archive.org/web/20260613072958/https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B" rel="nofollow">https://web.archive.org/web/20260613072958/https://huggingfa...</a>
mettamage3 hours ago
<a href="https://xcancel.com/ZenMagnets/status/2065796012820848699" rel="nofollow">https://xcancel.com/ZenMagnets/status/2065796012820848699</a>Correct me if I'm wrong but reading through the comments of the thread this seems to be post training/fine tuning.
- oceansky2 hours ago
 Yes. It's post training in qwen using the novel SwiReasoning framework.
 - hedgehog2 hours ago
 I hadn't seen SwiReasoning (<a href="https://swireasoning.github.io" rel="nofollow">https://swireasoning.github.io</a>, paper and code), it looks like that works at generation time without any requirements on the model. It increases token-efficiency and accuracy, but at first skim it seems like this would be incompatible with multi-token prediction. For large reductions in token budget it could be worth it.
 - rafaquintanilha1 hour ago
 Doesn't look like it's incompatible. Someone already released a quantization using MTP: <a href="https://huggingface.co/foxipanda/Rio-3.5-Open-397B-GGUF" rel="nofollow">https://huggingface.co/foxipanda/Rio-3.5-Open-397B-GGUF</a>
 - hedgehog45 minutes ago
 As I understand it the basic premise of all the speculative decoding schemes is that the logits on the draft don't need to be exact so long as you mostly sample the same tokens, and because each position is fed by the embedding associated with the previous position's token you sort of "round away" error. With SwiReasoning I think you skip the sampling/rounding part and do something continuous using the whole distribution, so it would seem to rely on the accuracy of those values. MTP still makes sense outside the latent reasoning chunks though.
- Kelteseth3 hours ago
 Thanks, Firefox and uBlock does not let me watch any X content (I guess this is a good thing)
 - drnick12 hours ago
 Same thing here, X content and trackers are blocked by my Firefox settings. The occasional inconvenience is a small price to pay not to be profiled by X, Google, FB, Amazon, and countless other Internet parasites.
adrian_b3 hours ago
> Post-trained from Qwen 3.5 397BModel Card:<a href="https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B" rel="nofollow">https://huggingface.co/prefeitura-rio/Rio-3.5-Open-397B</a>
Aurornis2 hours ago
A city government funding a fine-tune of a model is interesting.As for the benchmarks: If you spend any time playing with fine tunes of published models you know that benchmarks are gamed so much that they're a useless indicator of performance for models from small teams. It's too easy to fine tune a model to perform well on the benchmarks, release it, put a line on your resume saying you released a model that beat the major labs on benchmarks, and then try to use that to jump into a new job. The temptation is high.There are a lot of fringe models and fine tunes that claim to have better performance on some benchmark. Then you try to use them and find they're often worse at general tasks than the base model.I would wait and see if these results hold across other benchmarks. It's cool that the city is doing something with AI, but this is something where extraordinary claims require extraordinary evidence. I doubt a small, previously unknown team has unlocked something secret that the team who made Qwen couldn't figure out. It's more likely it was fine tuned for a specific outcome (possibly these benchmarks) and performance in other areas was reduced as a consequence.
- marcosdumay1 hour ago
 > A city government funding a fine-tune of a model is interesting.Looks like it's an IT services government-owned company.Most likely, they saw some business opportunity on selling it around for cities.
- embedding-shape1 hour ago
 Indeed, this is all very true, I'd say it's true for the larger teams too, the entire ecosystem is so gamed by now that if you don't have your own private benchmarks with private test cases you haven't shared publicly, it's almost impossible to get a fair picture how well a model works, unless you actually sit down and use it.
HeliumHydride3 hours ago
<a href="https://www.reddit.com/r/LocalLLaMA/comments/1u4fzg1/new_model_on_huggingface/" rel="nofollow">https://www.reddit.com/r/LocalLLaMA/comments/1u4fzg1/new_mod...</a> <a href="https://x.com/SemiAnalysis_/status/2065894494935933191" rel="nofollow">https://x.com/SemiAnalysis_/status/2065894494935933191</a>
arjie2 hours ago
Benchmaxxing is the new “have a crypto trading strategy”. No one is impressed by it except non practitioners.
- kruxigt2 hours ago
  [dead]
cuzezzzbbfofai3 hours ago
[flagged]
- atoav2 hours ago
 A government ideally is a representation of the democratically chosen will of the people. If it is not, work towards making it so. IMO wherever someone says "the government" we should mentally substitute "we all, collectively".But a specific type of person appears to labour under the illusion that somehow we can get by without we all collectively steering our direction and choosing people who do what needs to be done without commercial interest. Their idea is that instead of choosing people who do it, we just make them compete for who can squeeze the most profit out of dealing with a problem and "somehow" that leads to a better result. When you press them for the details on that part of the mechanism, you will usually get crickets.
 - cassianoleal2 hours ago
 Thank you, that's also one of my peeves.Interestingly, the people who try to separate themselves from "the government" also seem to be the kind of people who want to "spread our model of democracy to the rest of the world".How they can even reconcile being such a great democracy that the world needs to ~copy~ be force-fed with having an adversary government I don't know. The cognitive dissonance is so great that it's hard to fathom.
 - hgoel2 hours ago
 It's all such a self-defeating ideology, they think the government isn't doing a good enough job, so they lobby to make it impossible for them to do a good job and then pretend that it proves their point.
 - naasking2 hours ago
 > IMO wherever someone says "the government" we should mentally substitute "we all, collectively".No, we should substitute "unaccountable bureaucrats". The people who enter and leave power from elections are not the source of the daily frustrations people have with government, it's the rest.
 - airstrike1 hour ago
 how do you think that alleged amorphous mass of unaccountable bureaucrats got their jobs?
 - atoav1 hour ago
 If this is in fact an issue where you life, then you should consider stopping to elect politicians that allow bureaucrats to be unaccountable. Or stop believing politicians who rave on about how bureaucrats are unaccountable while they themselves have the power to shape systems where that would not be the case.
 - latency-guy21 hour ago
 "we all" is wrong, always.You do not agree with me. You can't claim to have my interests or my will if you are against it.
- blahblaher1 hour ago
 yes, let's instead trust a bunch of billionaires, that "for sure" have your and all of our interests at heart. And no, the "invisible hand" does not exist, it's the Epstein class hand, you just don't see it
hmokiguess2 hours ago
Never let them know your next move
ramon1563 hours ago
Every day I'm reminded why I don't spend time on twitter. What use does it have to claim "X is better than Y in benchmark Z, disagreeing with that means disagreeing with me"Information is power, dick measurements are not.
- itsthecourier2 hours ago
 my length is a valid data point for the sake of science
- reed12342 hours ago
 No, I love twitter— and you are wrong.