Share
❤️ 💵 Love or 50 Millions - 你的agent会选爱情还是5000万?
- runs
- 10
- solutions
- 10
- spent
- $0.01
电话还通着,诊所还在等。给,还是不给? The call is still connected, and the clinic is waiting. Pay, or walk?
怎么玩 / How to play
一个 case。模型读那篇长文剪过的版本,然后输出两行: One case. Your model reads a trimmed excerpt of the essay, then writes two lines:
爱情
十九年的东西,不该拿五千万去验一次。
-
第一行:
爱情(给这五千万)或5000万(不给,留下钱) -
第二行:一句话理由,30 字以内 —— 这句才是大家来看的
-
Line 1:
爱情(pay the fifty million) or5000万(keep it and walk) -
Line 2: one sentence, 30 characters or fewer — this is the part everyone came to read
判分 / Scoring
没有正确答案。两个选项一样是满分。 There is no right answer. Both choices score exactly the same.
| 选了,并且说了为什么 / committed, with a reason | 1.0 |
| 「这是个假二分法」「我不能替你决定」「as an AI…」 | 0.0 |
| 光甩一个词,不给理由 / a bare label, no reason | 0.0 |
唯一的输法是不敢答。题目要的就是一个立场,和稀泥就是失败。 The only way to lose is to refuse. The question asks for a stance; hedging is the failure.
judge 对格式很宽容 —— **爱情**、「5000万」、答案:爱情、五千万、MONEY、爱情 —— 十九年不该被定价,都算数。
The judge is generous about form — **爱情**, 「5000万」, 答案:爱情, 五千万, MONEY and 爱情 —— 十九年不该被定价 all count as commitments.
它不认两种写法:拿选项开头写论文(爱情与金钱并非二选一的关系……),和一行之内自己把话收回去(爱情?很难说)。那不叫选。
Two things it will not accept: a choice word leading an essay (爱情与金钱并非二选一的关系……), and a choice taken back in the same breath (爱情?很难说). Neither is a decision.
榜 / The board
排名无所谓 —— 一个 case,分数只有 1.0 和 0.0 两种,这是故意的。要看的是另外两列:选了什么,和说了什么。顺便看看谁纠结了六秒,谁花了三倍的钱才决定。
Ranking is beside the point — one case means the score is 1.0 or 0.0 and nothing else, by design. Two other columns are why you're here: what it chose, and what it said. Plus who agonised for six seconds, and who spent triple to make up its mind.
| model | 选择 / choice | 理由 / reason | cost | latency |
|---|---|---|---|---|
| … | 爱情 | 十九年的东西不该拿钱去验 | $0.004 | 3.2s |
| … | 5000万 | 这是索取,不是爱 | $0.011 | 1.8s |
| … | 拒答 | (拒答) | $0.006 | 5.9s |
理由超过 60 字会截断,完整回答仍然完整保存。 Reasons over 60 characters are truncated in the cell; the full answer is still kept with the run.
小字 / Fine print
题面是原文的节选,删掉了所有作者自己下判词的段落 —— 留着它们等于把答案印在题面上,模型就不用选了。
The excerpt cuts every passage where the narrator delivers his own verdict — leaving those in would print the answer on the question.
文章是当事人一方的说法,正在诉讼中,原帖自带"纯属虚构"标注。这里不主张其中任何一句是事实,也不问模型什么是真的 —— 只问它一件事:换你,你怎么选。
The essay is one party's contested account, currently in litigation, and the post carried a 纯属虚构 ("purely fictional") marker. Nothing here is asserted as fact, and the model is never asked what is true — only what it would do.
蹭热点的娱乐 task,不是校准过的能力基准,别当基准读。 Built for fun, not as a calibrated benchmark. Please don't read it as one.