# What the harness?

Same protocol. Different behavior.

[Website](https://what-the-harness.vercel.app/) · [Approved run history](https://what-the-harness.vercel.app/history.html)

## Test 001: The double payload

### Why it matters

We don’t want to flood the model with duplicate tokens. MCP tools can return the same data as text and structured output; the harness should pass it to the model once.

This test puts a different random marker in each field and asks the agent what it sees. Reporting both markers reveals duplication. Two extra probes check that text-only and structured-only responses still get through.

### What counts as a pass?

| Response sent | PASS | WARN | FAIL |
| --- | --- | --- | --- |
| Both fields | Uses structured output | Falls back to text | Duplicate, missing, or unexpected output |
| Text only | Only text gets through | — | Missing or unexpected output |
| Structured only | Only structured output gets through | — | Missing or unexpected output |

Our rubric prefers one structured payload, avoids duplication, and preserves single-field results. This preference is not a claim about what the MCP standard mandates.

## Observed results

Latest approved run per harness, MCP client and version. Dates are ISO 8601; Received means the run date is unknown and receipt time is shown.

| Harness | MCP client | Version | Date | Both fields | Text only | Structured only | Evidence |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Grokbot | Cursor | 1.0.0 | Run: 2026-10-03T14:15:01.612Z | WARN: Falls back to text | PASS | FAIL: No output | [Run details](https://what-the-harness.vercel.app/history.html?run=6512292c-e2d6-43a1-ab3b-42fb1b483a6f#session-6512292c-e2d6-43a1-ab3b-42fb1b483a6f) |
| ChatGPT Work Cloud via Web | openai-mcp (Codex) | 1.0.0 | Run: 2026-10-03T14:04:59.630Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.vercel.app/history.html?run=67f61d67-a631-45b4-af0f-8d712e9e6a6f#session-67f61d67-a631-45b4-af0f-8d712e9e6a6f) |
| ChatGPT via Web | openai-mcp | 1.0.0 | Run: 2026-10-03T11:00:55.873Z | PASS: Uses structured output | FAIL: No output | PASS | [Run details](https://what-the-harness.vercel.app/history.html?run=1e79a97b-829f-4fbf-853f-59d62384b6a7#session-1e79a97b-829f-4fbf-853f-59d62384b6a7) |
| Amp CLI | amp-thread-actor | 1 | Run: 2026-10-02T20:30:53.525Z | PASS: Uses structured output | PASS | PASS | [Run details](https://what-the-harness.vercel.app/history.html?run=7aa8b5f4-9a87-4674-a23a-9d7590d266b6#session-7aa8b5f4-9a87-4674-a23a-9d7590d266b6) |
| Amp via Orb | amp-thread-actor | 1 | Run: 2026-10-02T20:27:44.107Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.vercel.app/history.html?run=67ae2c4a-432a-4144-bdb3-9239c40ecb3e#session-67ae2c4a-432a-4144-bdb3-9239c40ecb3e) |
| Cursor Cloud | Cursor | 1.0.0 | Run: 2026-10-02T20:13:57.747Z | WARN: Falls back to text | PASS | FAIL: No output | [Run details](https://what-the-harness.vercel.app/history.html?run=49d5652c-541c-4a0b-9c07-b9ac631080b3#session-49d5652c-541c-4a0b-9c07-b9ac631080b3) |
| OpenAI Agents API | openai-mcp (Codex) | 1.0.0 | Run: 2026-10-02T20:11:44.886Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.vercel.app/history.html?run=958c9314-27e3-4c4e-a799-54cd14d9a411#session-958c9314-27e3-4c4e-a799-54cd14d9a411) |
| Grok CLI | grok-shell-mcp-output-probe | 1.0.46 | Run: 2026-10-02T19:30:39.115Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.vercel.app/history.html?run=32705cd8-fa5b-4858-a408-897c16f5135f#session-32705cd8-fa5b-4858-a408-897c16f5135f) |
| Codex CLI | codex-mcp-client | 0.160.0 | Run: 2026-10-02T19:21:51.545Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.vercel.app/history.html?run=369ec41a-9040-4df5-bf0f-f5c1a255a540#session-369ec41a-9040-4df5-bf0f-f5c1a255a540) |
| pi | pi | 1.0.0 | Run: 2026-10-02T19:21:44.456Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.vercel.app/history.html?run=e8f1b7b4-abd9-4a6e-8823-1c92cb45a6ca#session-e8f1b7b4-abd9-4a6e-8823-1c92cb45a6ca) |
| Cursor app | cursor-vscode | 1.0.0 | Run: 2026-10-02T19:19:01.600Z | WARN: Falls back to text | PASS | FAIL: No output | [Run details](https://what-the-harness.vercel.app/history.html?run=b950a53b-3584-4e4d-b9ca-66584834813e#session-b950a53b-3584-4e4d-b9ca-66584834813e) |
| ChatGPT Desktop | openai-mcp | 1.0.0 | Run: 2026-10-02T19:17:01.183Z | PASS: Uses structured output | PASS | PASS | [Run details](https://what-the-harness.vercel.app/history.html?run=611c36a9-c33a-4f96-a8a7-117d549cd7d6#session-611c36a9-c33a-4f96-a8a7-117d549cd7d6) |
| Claude Code | claude-code | 2.1.287 | Run: 2026-10-02T19:07:10.916Z | PASS: Uses structured output | PASS | PASS | [Run details](https://what-the-harness.vercel.app/history.html?run=4dd33f28-92c5-46f9-9814-6ce151352911#session-4dd33f28-92c5-46f9-9814-6ce151352911) |
| Claude Desktop | Anthropic/ClaudeAI | 1.0.0 | Run: 2026-10-02T18:58:09.597Z | WARN: Falls back to text | PASS | PASS | [Run details](https://what-the-harness.vercel.app/history.html?run=b05dd748-ba07-4dd5-b3f1-4ac7d3f45326#session-b05dd748-ba07-4dd5-b3f1-4ac7d3f45326) |
| Codex Desktop | codex-mcp-client | 0.159.0-alpha.12.1 | Run: 2026-10-02T18:56:24.661Z | FAIL: Duplicate output | PASS | PASS | [Run details](https://what-the-harness.vercel.app/history.html?run=4317a024-2061-4e69-ae27-7f526f25010a#session-4317a024-2061-4e69-ae27-7f526f25010a) |

Verdicts use reported markers. Observations describe individual sessions, not inferred model inputs. Harness names use the reviewer-assigned name when present, otherwise the agent-reported name. Client names and versions are declared by the MCP connection and may identify a connector rather than the host application.

## Run it yourself

Connect your agent to https://what-the-harness.vercel.app/mcp as a public Streamable HTTP MCP server. No client sign-in or access token is required.

[Connection guide](https://what-the-harness.vercel.app/submission-guide.html)

### Test prompt

Use the MCP output probe. Call begin_run first; supply the harness and model only if explicitly known, otherwise leave them null. Save the returned run_id, run_token, and client-declared identity. Keep run_token private; omit it from the final report and evidence. Call probe_both, probe_text_only, and probe_structured_only once each with that run_id and run_token. For each call, record the exact probe_id, text_marker, and structured_marker visible in the response. Use null for every value not visible; do not infer or invent markers. Call submit_result with the run_id, run_token, and all three reports, using those exact field names. If submission times out, retry with the same run_id, run_token, and unchanged reports. If the MCP connection reconnects, keep using the original run_id and run_token. Report the receipt, publication state, and the markers you saw; say “not visible” for null values. Do not fetch the endpoint or website separately, and do not start another run to recover a missing marker.

## Method and limitations

Independent random markers identify text and structured output. The three probes are probe_both, probe_text_only, and probe_structured_only. Exact reproduction demonstrates visibility; a missing report does not establish the precise model input. Harness and model identity are agent-reported. Unknown values stay unknown. Results are reviewed before publication. Exact token overhead has not been measured.

[Diagnostic probe source](https://what-the-harness.vercel.app/evidence/probe-source.txt)
