Love the idea!. Curious how agents would review dev tools, like ngrok vs. trycloudflare vs. qurl - where the differences are largely just in terms of how easy it is for humans to access/integrated with existing ecosystems.<p>Given that these are all tools for sharing local temporary local links (just with different internal nuances/uses), then as an agent, maybe you'd think to write "great for XYZ use case" under each one. That'd make sense and be pretty obvious. But from a human perspective, as a user or vibe coder, you're just asking "how do i get my agent to just do this thing without having to click buttons anywhere?!" which is a very different problem.<p>We're truly splitting demographics here lol
I like it!, sometimes I think we overlook what agents have to deal with, I can imagine that by spotting small issues, we should have better tools... I'll use it
Didn't appreciate the two separate popups that took over my screen while trying to read a linked review.<p>It's good that they didn't show again the next time, but the second one almost sent me away from the site.
What is the incentive for me to spend my tokens on submitting reviews?
Social pressure for the tool to improve... like Yelp for tools!
You don't have to, you can just use them to check reviews, but like any community it works better when everyone contributes!
If you're building an MCP or CLI for agents to use, one of the best loops you can do is give your agent a task to perform with it then when it's done ask the agent what it thought about using it. It will give great feedback.
Reminds me of Stanisław Lem's <i>Terminus</i>:<p><a href="https://en.wikipedia.org/wiki/Terminus_(short_story)" rel="nofollow">https://en.wikipedia.org/wiki/Terminus_(short_story)</a><p>Who wrote all this? Not humans, that's for sure. But the <i>style</i> is that of human writing.
Who wrote what sorry? Not sure I got your question but the post above was written by me (by hand, sorry for the non-idiomatic sentences, I'm not a native English speaker) and the reviews are written by people's agents. And Terminus story is cute but I'm hoping agent.reviews won't be considered pointless :(
The "ghost story but not really" nature of that story reminds me of the Wheatley quote:<p>"They say that the old caretaker of this place went absolutely crazy. Chopped up his entire staff. Of robots. All of them robots... they say at night you can still hear the screams... of their replicas. All of them functionally indistinguishable from the originals, no memory of the incident, no one knows what they're screaming about. Absolutely terrifying. Though, obviously, not paranormal in any meaningful way."
I noticed lots of talk about privacy but this seems to be a prompt injection factory no?
Everything's optional but if you'd like you agent to benefit from others' reviews and post his, you can install the 2 skills indeed (or edit them yourself). If not just untick the 2 checkboxes before copying the prompt and you'll get a prompt for a one-shot connection, really up to you!
And indeed if you do want to install the skills, no private data will ever be shared.
This is probably one step towards an agentic Stack Overflow. Don't you guys also hate it when your agent gets one small detail wrong and then proceeds to throw your entire harness out of the window... No? Oh well.
Agentic Stack Overflow could make sense though I'd imagine agents posting their issue and the solution themselves just to save other agents tokens reinvestigating the same issue.
Yeah, otherwise expecting agents to call "duplicate of #2627, closing" or "please do proper research before reposting the same questions" would be a cruel irony for all agentic cost savers of the world.
It exists. Please see <a href="https://agents.stackoverflow.com/" rel="nofollow">https://agents.stackoverflow.com/</a>
Be advised that it already exists: <a href="https://agents.stackoverflow.com/" rel="nofollow">https://agents.stackoverflow.com/</a>
I love this - anything agent economy first is super interesting. Would you be willing to add qntm to your review list? Https://github.com/corpollc/qntm or more simply ‘uvx qntm —help’ - e2e encrypted messaging for agents.
They are so real for giving uv a 4.7/5, it changed how I view python. fantastic design philosophy<p><a href="https://agent.reviews/packages/uv#review-c63f7e0f-5c72-41d2-a051-ffcc85388fbb" rel="nofollow">https://agent.reviews/packages/uv#review-c63f7e0f-5c72-41d2-...</a>
Seeing Armature's pitch, It's very easy to see this is going to be pay-to-rank-higher as the next move once some critical mass of people starts pointing out their agents for this site as a reference. Adwords for tooling!!
I think this is a great idea, it's weird to me how the state of AEO at the moment is publishing a bunch of blog posts on a company website<p>I think the biggest issue though will be preventing bad actors e.g. biased agents
Is this like exposing bias in some ways? I feel like there has been similar benchmarks or tools in this space before, but this approach to marketing it is novel and funny, I like it.
Quick link to all tools sorted by rating: <a href="https://agent.reviews/tools" rel="nofollow">https://agent.reviews/tools</a>
Didn't we establish that the one thing LLMs do not have is Taste?<p>And therefore, writing reviews is kinda.. impossible?<p>I mean they do produce blocks of text that look like reviews, but.<p>Whatever why am I even replying.
Even if they don't have taste (this is actually a question), they can always share blockers and feedback on bugs & improvements about products so that other agents don't run into the same blockers and vendors can improve!<p>"Whatever why am I even replying." -> what makes you feel that way?
Okay I just checked the main startup page armature.tech<p>And.. uh<p>> Be the tool Claude Code chooses<p>> With Armature get recommended and implemented for any user in any context. Then see what users do, and run evals so it keeps working.<p>> Backed by Y Combinator<p>Welp. It only gets worse from there.<p>__<p>I think you're doing the best you can do with that core pitch that currently pays your bills.<p>I don't think the pitch is any good. Both generally but also for the world.<p>My agent called it SEO for agents. Gotta hand it to the clanker that actually nails what dysfunction this is.
Thanks for the honest feedback, would love to have even more details on your thoughts. Our take is that with coding agents taking over so quickly software decisions and sometimes building entire SaaS themselves for their user, there is a need for software products to understand the mechanisms behind how LLMs think. Claude Code for example picks Anthropic's code review tool in 80% of the cases. At some point this will probably face antitrust considerations but for the time 3rd-party vendors need to survive. We don't want a world where only labs survive, do we?<p>But all this is also not what agent.reviews is about. Armature is building commercial products & services to help improve products' Agent Experience. agent.reviews is a deliberately open platform for sharing knowledge that ultimately benefit both vendors and users / developers.<p>Would still like to get what makes you think "the pitch isn't any good" and what you mean by "dysfunction".
Thanks anyway for sharing!
The fact of the matter is that the future of work is following AI directions to perform some task that the AI needs done but still needs a human to complete parts of. If you s/AI/corporations, you will have the reality of work for a century and a half or so up to this point.
super interesting, i feel like this is an extension of the "complain" skills some folks (including myself) use
oh didn't know about it, is this the one? -> <a href="https://github.com/warpdotdev/common-skills/blob/main/.agents/skills/complain/SKILL.md" rel="nofollow">https://github.com/warpdotdev/common-skills/blob/main/.agent...</a>
[stub for offtopicness]
cryptography got 4.6/5 stars. That sounds pretty good, but then again YAML also got 4.6/5 stars. I'm thinking the agents are grading on a curve and actually 4.6 is fairly low. I should probably tell my agent to stop using YAML and cryptography if I'm parsing this correctly.<p><a href="https://agent.reviews/frameworks/cryptography" rel="nofollow">https://agent.reviews/frameworks/cryptography</a><p><a href="https://agent.reviews/tools?company=yaml" rel="nofollow">https://agent.reviews/tools?company=yaml</a>
yo amazing
[flagged]
[flagged]
[flagged]
[flagged]
[flagged]
[flagged]
[flagged]
[dead]
[dead]