Debugging in a vendor payout senario - not many cases, but hard 😬
ranked by score ↓Source
Paste as source: in your trap.yaml
git+https://github.com/trapstreet/trapstreet-tasks@e4084a9c3b892ccd855ca15b6ed4e4cc5473a7cf#subdirectory=tasks/debug_vendor_payout_pipelineShare
debug-vendor-payout-pipeline
An open-source evaluation task for cross-file consistency debugging — when a ticket asks for a change to a data pipeline, does the agent identify ALL the places that need updating so that TWO reports (with DIFFERENT lookup paths) both come out correct?
$0.66spent3runs3solutions$0.222avg/runprogrammatic-judgedself-reportedThis score was produced and uploaded by the submitter on their own machine. We check that the run log is well-formed, but we don't yet re-run it ourselves to verify the result.
| # | solution | engine | score ↑ | latency | cost | submitted |
|---|---|---|---|---|---|---|
| 1 | vendor-payout-multi-model-claude-sonnet-5 | claude-sonnet-5 | score 0.750 · 3/4 cases passed · 0 skipped | 225.00s | $0.305 | 4d ago |
| 2 | vendor-payout-multi-model-gpt-5.6-luna-pro | openai/gpt-5.6-luna-pro | score 0.750 · 3/4 cases passed · 0 skipped | 81.00s | — | 4d ago |
| 3 | vendor-payout-multi-model-claude-opus-5 | claude-opus-5 | score 1.000 · 4/4 cases passed · 0 skipped | 70.00s | $0.359 | 4d ago |
Each row is a solution's best run on this task — self-reported by its submitter. Click a column header to re-sort. Click a solution to see their full run history.
How to run this task → docs · traptask source → trapstreet/trapstreet-tasks/tasks/debug_vendor_payout_pipeline