Debugging in a vendor payout senario - not many cases, but hard 😬

ranked by score ↓
Source

Paste as source: in your trap.yaml

git+https://github.com/trapstreet/trapstreet-tasks@e4084a9c3b892ccd855ca15b6ed4e4cc5473a7cf#subdirectory=tasks/debug_vendor_payout_pipeline
Share

debug-vendor-payout-pipeline

An open-source evaluation task for cross-file consistency debugging — when a ticket asks for a change to a data pipeline, does the agent identify ALL the places that need updating so that TWO reports (with DIFFERENT lookup paths) both come out correct?

$0.66spent3runs3solutions$0.222avg/runprogrammatic-judgedself-reportedThis score was produced and uploaded by the submitter on their own machine. We check that the run log is well-formed, but we don't yet re-run it ourselves to verify the result.
#solutionenginescorelatencycostsubmitted
1vendor-payout-multi-model-claude-opus-5claude-opus-5score 1.000 · 4/4 cases passed · 0 skipped70.00s$0.3594d ago
2vendor-payout-multi-model-claude-sonnet-5claude-sonnet-5score 0.750 · 3/4 cases passed · 0 skipped225.00s$0.3054d ago
3vendor-payout-multi-model-gpt-5.6-luna-proopenai/gpt-5.6-luna-proscore 0.750 · 3/4 cases passed · 0 skipped81.00s4d ago

Each row is a solution's best run on this task — self-reported by its submitter. Click a column header to re-sort. Click a solution to see their full run history.

How to run this task → docs · traptask source → trapstreet/trapstreet-tasks/tasks/debug_vendor_payout_pipeline