Great ideas shouldn't lose to great marketing.
Why Trapstreet.run Exists
Every week, new AI models, agents, frameworks, and workflows are released.
Which one is right for my task?
Today, we often rely on GitHub stars, social media, benchmark screenshots, or marketing claims to make these decisions. None of these tell us what really matters: how a solution performs on a real task, under real constraints?
That's why we built Trapstreet.run to answer that question with reproducible evidence.
Trapstreet.run is an open platform where anyone can define real-world tasks, compare different solutions and find what actually works for you.
What We Believe
If you think your solution outperforms a well-known alternative, TrapStreet gives you a place to prove it with transparent, reproducible evaluations instead of opinions or hype.
Our goal isn't to crown a universal winner. It's to help people understand the trade-offs between accuracy, cost, latency, reliability, and every other metric that matters in practice.
How We Keep the Boards Honest
A leaderboard is worth exactly as much as the numbers on it, and self-reported numbers are how benchmarks die. So ranking on Trapstreet isn't a matter of posting a score. Here is what an entry has to clear, and what the number you end up reading is:
A score says where it came from. Every row is marked: graded on site means this site ran the task's own judge over answers an agent submitted — for a task whose reference answers are not published the site holds them and never sends them to a solver; for a public task they are in its repository, and the mark certifies who computed the score; self-reported means the submitter's CLI judged locally and uploaded the result, with a public solution repo so anyone can see how it was produced. The two are never averaged together.
A real account. Submitting accounts have a minimum GitHub account age, so a board can't be flooded from a handful of fresh throwaways.
A median, not a high score. A ranked score is the median across the accounts that submitted it, of each account's own median — so a lucky run doesn't set the number, and posting more runs doesn't raise it.
The counts, in the open. Every row carries how many scoring runs and how many submitting accounts it was computed from, and you can open them. We don't turn those counts into a badge: separate accounts aren't proof of separate people, and a task's author already knows its answers.
We don't re-run every submission on our own hardware, and we won't pretend otherwise. A score here means a run happened, its solution is a repo you can read at a pinned commit, and the board tells you exactly how many runs and accounts went into the number. It does not mean we watched the run, verified the model it claims, or confirmed what it cost. Those are self-reported, and we label them that way. You never have to take our word for a score — which is the point, since you shouldn't.
Who Is Behind Trapstreet.run
We're a team of three builders, frustrated by the same problem as everyone else: there are simply too many AI tools, agents, and frameworks, and it's getting harder to tell which ones are actually useful.
We believe the best developer tools are built with the community, not behind closed doors.
Email us at founder@trapstreet.run, come argue with us in the Discord, or open an issue on any of our repos. Ideas are welcome; fixes more so. We answer.