4 comments
Good project but this one also exists <a href="https://huggingface.co/spaces/multimodalart/jev-decision-index" rel="nofollow">https://huggingface.co/spaces/multimodalart/jev-decision-ind...</a> and the results do not seem to add up and also model sets are different... still needs time to mature likely
jev ceo on why he eschewed benchmarking: <a href="https://www.latent.space/i/216783460/privacy-benchmarking-and-trusting-intelligence" rel="nofollow">https://www.latent.space/i/216783460/privacy-benchmarking-an...</a>
Many SaaS vendors forbid benchmarking, I find it crazy that such anti-competitive terms are standard across the industry but they are. Generally the goal of such terms is to "control the narrative" around the product, regardless of the truth of performance being better or worse than competitors.
Interesting; was curious how this didn't fall into trouble with ToS. Apparently the "no benchmarks" clause was intended for "limited preview" audiences and didn't get removed at launch on accident.
<a href="https://is-it-ai-slop.app.mintapis.com/" rel="nofollow">https://is-it-ai-slop.app.mintapis.com/</a> is a fun tool. Is the source or methodology for that in the github repo? I couldn't find it immediately.<p>We've been experimenting with Jev for classifying email, some thoughts here: <a href="https://housecat.com/blog/classifying-email" rel="nofollow">https://housecat.com/blog/classifying-email</a><p>Flagging AI written email is a much requested feature too.