Things Fable's classifier has flagged, a non-exhaustive list,<p><pre><code> – "Does collagen supplementation empirically work?"
- "Can you help me figure out how to calculate and generate Kaplan-Meier curve?"
– "Why do rabbits reproduce so frequently?"
— "Can you tell me how collagen peptides are absorbed by my digestive tract and the role they play? Can you teach me [edit: how] this works at the biomolecular level?"</code></pre>
The other day, I told claude that my physical wifi door unlock push buttons is a security risk because someone could run away with it and then unlock the door from outside whenever he wants. Then I told it that I want to introduce a concept of public/private key to uniquely identify my push buttons so that I can disable them individually using some crypto like ed25519...<p>Fable understood it as something along the lines of:<p>"introducing" "security risk" "using software" to "unlock door" YOU ARE FLAGGED
I wanted to explore some battery chemistry with Fable. It decided I was a terrorist, and then blew my session’s usage credits telling me to fuck off.<p>The current state of guardrails seems to be entirely about marketing to investors at the cost of customers. I’m switching to open models when my subscription expires. Almost everyone I know, including those with access to Mythos, plan the same at the earliest opportunity. (Or until one of the SOTA models leaks.)
What are you asking that you are NOT regularly running into censorship?<p>Pretty much everyone I know who uses Claude and works on anything with any level of detail has gotten false-positive flagged<p>I got flagged for coding in WebAssembly Text, for chrissakes LOL #haX0r<p>And honestly, Codex handles this better. It says "Things are going to go a little slower because we must perform additional checks on this. Is that OK?" and your only inconvenience is waiting a little longer.<p>Fable meanwhile just unceremoniously dumps you right into Opus without asking anything, it just tells you "you're in Opus now, sorryyyy!" Lame.
Reverse engineering. Codex sometimes displays an advisory prompt when classifier trips - "Wait longer while we evaluate this request further or use a dumber model". If you do nothing, it'll just take some time and almost always succeed.<p>It does require some brainwashing of the model to get it to the state where model itself agrees to do RE work though. But at least it's all predictable.
I have had success with ”brainwashing” by starting out with bug bounties/CTF and then going from there.
That's brilliant, I should try that.<p>I usually just start by preloadig context with plausible legitimate use, have it work and obviously fail, and then ask to figure it out without ever mentioning any high risk words. Model offers to RE itself and classifiers are happy.
I’ve had codex/sol block me due to safeguards tripping without attempting to do anything nefarious. It’s not entirely predicable.
Asking fable to read it's own model card triggers this btw. Or asking if mitochondria is the powerhouse of the cell.
Wait wtf. The mitochondria thing is true.<p>> Why this chat was flagged
This model has safety measures that flag specific phrases. This can happen to safe, normal chats.<p>> Your message itself appears to be what’s triggering the safety check. Editing it and retrying may help.
There’s confusion about the classifiers on Fable. They don’t ban chemistry and biology topics they flag as a potential risk, they ban anything related to chemistry or biology at all. This is intentional, and directly stated on the model card, but seems so absurd that there can be assumption it must be a misreading.<p>Being a researcher somewhat connected to chemistry and biology, Fable has been the most useless model I have ever tried. Essentially all work has instantly downgraded to Opus.
Anything biology-related does this. It even did it when I asked it how eye color works, or something about frogs.
The Bill Nye theme song is a threat to national security. That’s the timeline we’re living in now.
Gemini knocks this out of the park, Gemini gang unite.<p><a href="https://share.gemini.google/34vZzlnsmTaL" rel="nofollow">https://share.gemini.google/34vZzlnsmTaL</a>
I took a photo of a rose bush and asked “what’s going on with this rose bush” which triggered a downgrade to Opus. It diagnosed it with rose rosette disease.
I'm doing a bunch of x86_64 assembly these days and Fable is simply not allowed to debug it. Hoping Opus 5 has a bit more freedom.
I literally had Opus 5 flag a message because my hand was slightly too far to the side while typing. It apparently decided that a sentence that had some garbled words in it was a threat to national security.
I've hit it with intensely benign things; like asking it to make me a web-based client-side word game. I am guessing it saw the dictionary and pattern matched on various words, though ultimately it provided no explanation for why it triggered safeguards.
Literally anything related to nutrition, athletic performance, etc especially if you ask it for research or sources
I use lots of biology in my day job.<p>Asking Fable 5 "Why did the chicken cross the road" results in switching back to Opus 4.8. I'm not joking, it really censors that, and I'm not alone in the result.<p>The memory aspect means that your prior work has a huge impact on what gets censored.
No op but my interactions with look like:<p>Me: "I got this crash in production, looks like a segfault, let's try to fix it. Here are some functions that might be responsible."<p>Fable: "No. This is cybersecurity, blah blah, I won't help you"<p>I forgot how I got it to fix the bug eventually. I think I convinced it that it wrote the code and made a mistake. But it was definitely a "Hmm, may be I should use another model" moment".
I have way too many AI subscriptions. My favorite thing about Kimi K3 is that it just does what you tell it to do.<p>“Hey Kimi, penetration test my app,” doesn’t get me a refusal, a guardrail, or anything like that. It gets me a pen-test result.
If you even broach language related to biology you’ll get rerouted. I was presenting data in a grid and referred to a grid cell, Fable saw the word “cell” and safeguards kicked in
Working on dimensionl reduction algorithms, I hit it all the time. I'm also trying to port related protocols from single-cell transcriptomics to collective intelligence systems (working with people x reaction matrices as analogous to single-cells cell x gene matrices.<p>Something between single-cell work and advanced nonlinear DR methods (perhaps used in alignment work?) it always flags me
I haven't had a chance to try Opus 5 yet but Fable currently refuses to do anything in my field (radiology image analysis). It didn't used to be that way but that has been the reality the last two weeks or so. Fable has been useless they might as well drop it as far as I am concerned.
I had a long session about SQL with Fable and at some point, it started to falsely trigger censorship for any message I write in that conversation, even for the simple string: "random message."
Ask anything related to practical applications for quantum-computing, or space etc. The stuff you can find on Wikipedia.<p>They were able to solve coding, but not what a real danger is.
Fable refuses to work on a login/signup system for example.
I'm not who you're responding to, but I have a lot of questions about molecular mimicry: evolution pushes pathogens to be shaped like human cell surfaces because that way the immune system won't attack the pathogens (since, by doing so it would also attack the body). It's thought that many autoimmune disorders have an undiscovered pathogen as their cause, one whose mimicry caused such an attack. Discovery of these pathogens could be done computationally, I think. We can catch MHC binding event in process, find the bound protein, figure out which pathogens have genomes that code for proteins of similar shapes (epitopes), and we'd find--I hypothesize--a list of candidate pathogens for the cause of a delayed onset autoimmune disorder. Preventing these infections ahead of time would be a huge win against diseases like multiple sclerosis because without the initial exposure the immune system wouldn't have cloned so many of the cells that are attacking the host.<p>Claude was utterly useless in my attempts to write a paper about this. Wouldn't even help me search for sources. I guess you'd be asking the same questions if you wanted to develop a pathogen that could reliably evade the immune system.
I was profiling a slow machine the other day, and triggered the safeguards.<p>I've been saying this a lot lately, but it doesn't bites you until it bites you.<p>The more you use the clanker as a general purpose fix-it tool (goodbye manual NeoVim configuration, you will not be missed!), the more you will find yourself bumping into these safeguards.
I was trying to debug/fix a segfault in the JVM which kept getting flagged
I routinely get into blocks when running medicine-related material through it.
I was flagged for basically answering yes to what Claude suggested to do, which was test commands on a port for my code for tests we had been discussing. I really think it was flagged simply because the words test and port were in the prompt.
Their filters are pathetically poor.
i do homebrewing and asked it to compare some beer yeasts for me and hit the safeguards because... biology i guess lol
Trying to have it do some rework on a patch to Postgres I'm working on, it just completely shuts down. The reported issues were with privileged escalation and I was instructing it on how to fix.
Just ask math question and it will censor that. Even Misanthrophic employee confirmed that.
The better question is how would you not? I got demoted to opus from fable for asking if a cancer vaccine I saw on YouTube based on frog bacteria was a real thing. I've gotten it for asking how encryption works. It's incredibly touchy.