Skip to main content
CONEX
Agent Smartness app icon

Agent Smartness

macOS 14 Sonoma or later

Track whether your AI coding agents are still as good as they were.

Download on the App Store (coming soon)

Price: Free, with every feature available. Optional “Coffee” tip as a consumable in-app purchase.

What is Agent Smartness

Agent Smartness is a macOS menu bar app that measures OpenAI’s Codex and your own installed Claude Code over time — using your own accounts and your own provider quota — and watches for change.

  • It measures the difficulty your agent can still handle, not a pass rate.
  • Your measurement history stays on your Mac. Nothing is sent to the developer.
  • It does not rank Codex against Claude Code. The comparison that matters is an agent against its own earlier measurements.

TermsSequential test: a statistical method that updates its verdict as each observation arrives, deciding without waiting for a fixed number of results. Consumable: an in-app purchase that unlocks nothing and can be bought again.

What it does

  • Measures the difficulty your agent can still handle — each question gets harder when the agent gets it right, and questions keep coming from near where it fails. The pass rate therefore stays near half, and what moves is how far the agent can go. Fixed questions get answered nearly every time: Codex scored full marks on 30 of our 30 measured runs, which spends tokens and measures nothing.
  • Two kinds of question — counting occurrences of a two-character pattern in a long string, and multiplying two large integers exactly. Each has one correct answer, and grading happens on your Mac — no second model is asked for an opinion.
  • The characters it can count through, the digits it can multiply — shown with the range the estimate places each in. That range narrows as measurements accumulate.
  • Change detection — tells you when the agent starts missing items at a difficulty it used to handle, using a sequential test that decides as the evidence arrives rather than waiting for a fixed window.
  • Value and format graded separately — showing the working is not a wrong answer, and only questions about format require an exact match.
  • Agentic task (opt-in, Claude Code only) — once a day, a small read-only tool-use task inside a temporary folder Agent Smartness creates itself. Codex does not currently support this axis.
  • Failures and outages recorded separately — timeouts, provider failures and quota exhaustion are never scored as a zero.
  • Missed items kept on your Mac — with the expected and given answers, so you can judge the grading yourself.
  • Token usage — sums the token counts your own Claude Code and Codex already record for their own sessions. The conversations are never read.
  • Demo mode — synthetic data, no network, no account. Try the full dashboard before connecting anything real.
  • And more — history and diagnostics with JSON export, a menu bar status view, Japanese and English UI, keyboard shortcuts, and VoiceOver labels.

How scores work

Scores are relative to each agent’s own baseline over time. Agent Smartness does not rank Codex against Claude Code, or claim that one is “better” than the other. The comparison that matters is an agent against its own earlier measurements.

Pricing

Agent Smartness is free, and every feature is available to everyone. An optional one-time “Coffee” tip (a consumable in-app purchase) is available for anyone who wants to support development — purchasing it unlocks nothing.

Privacy

Agent Smartness has no analytics and sends nothing to its developer. Measurement history is stored in a local SQLite database on your Mac.

When a measurement runs, the only thing sent to the provider you chose (OpenAI or Anthropic) is the benchmark question text Agent Smartness generated — for the Agentic task, the small files it creates as well. Your source code, files and credentials are never sent.

For the full details, see the Privacy Policy.

Requirements

  • OS: macOS 14 Sonoma or later
  • Accounts: your own OpenAI (Codex / ChatGPT) sign-in, and/or your own separately installed Claude Code with its own plan. Neither is required to try Demo mode.
  • Quota: measurements consume your own provider quota.
  • Interface languages: Japanese / English

Contact

For questions or requests about Agent Smartness, please use the contact form.

Not affiliated

Agent Smartness is not affiliated with, endorsed by, or sponsored by OpenAI or Anthropic. All trademarks belong to their respective owners.