Claude
Claude is the reasoning layer. It is not the source of truth.
Coordinating people means handling language that resists rigid forms and rules. That is the part Claude is for. Everything that must be correct is kept out of its hands.
Integration status
Exactly what exists.
- Role in the design
- Claude is the reasoning and orchestration layer for the parts of coordination that resist rigid rules.
- Implemented Prototype (Built in the repository. Not used with real players.)
- A server-side intent-extraction endpoint (POST /api/intent) using the official Anthropic SDK. It sends one player request to Claude with a single forced tool, then validates the result against a strict schema before anything else can see it.
- Status
- Integration boundary implemented and unit-tested against fixtures. Not connected to real players. No production usage is claimed.
- Live demo
- The public live free-text demo is disabled unless the deployment is configured with credentials. The worked example on /technology is deterministic and makes no model call.
- Model choice
- Model identifier is read from the ANTHROPIC_MODEL environment variable; it is never hard-coded or client-controlled.
01Why Claude
Because the hard inputs are sentences, not fields.
A coordination request mixes a firm limit, a flexible wish, a conditional and an unstated assumption in one line. Turning that into something a program can use requires reading comprehension, judgement about what is firm, and knowing when to ask rather than assume. Those are the capabilities Doubles depends on Claude for. Product quality should rise as the model gets better at them, and the architecture is built so that it can adopt a better model without loosening any hard rule.
02Where Claude sits
The judgement calls.
- Intent extraction. Natural-language request in, structured intent out.
- Hard versus soft. Deciding which stated conditions are non-negotiable.
- Missing information. Identifying what must be asked, and when inference is safe enough.
- Group proposals. Suggesting configurations across incomplete groups and uncertain skill. Research (A question we are working on. No result is claimed.)
- Tool selection. Deciding which scoped tool to request next. Planned (Intended. Not built.)
- Exceptions. Cancellations, unusual requests, contradictory instructions. Research (A question we are working on. No result is claimed.)
- Communication. Drafting messages to players that reflect the situation. Planned (Intended. Not built.)
- Longer workflows. Carrying a coordination effort across hours or days without losing the thread. Research (A question we are working on. No result is claimed.)
03Where Claude does not sit
Anything that must be true, or can’t be undone.
Claude reasons. Deterministic systems decide what is true.
A model may propose a group. It cannot override an unavailable court.
Claude: model reasoning
- Natural-language intent understanding
- Separating hard constraints from soft preferences
- Reasoning across incomplete groups and uncertain skill
- Choosing which scoped tool to request
- Handling cancellations and unusual exceptions
- Drafting context-aware messages to players
Software: deterministic state
- Player identity and authoritative state
- Court capacity and venue inventory
- Final availability and scheduling conflicts
- Booking state and payment state
- Exactly-four validation before any action
- Irreversible actions and audit logging
A model may propose a group. It cannot override an unavailable court, mark a payment as received, or confirm a booking. Those facts come from authoritative systems, and the validator refuses any proposal that disagrees with them.
04The implemented boundary
A small, real proof of architecture.
The repository contains a server-side endpoint that converts a free-text request into a validated intent. It was built to make the boundary concrete, not to simulate usage.
Bound the input
Control characters are stripped, length is capped at 500 characters, and the text is wrapped in delimiters and labelled as untrusted data.
Force a single tool
Claude is given one tool,
record_player_intent, whose input schema is generated from the validation schema. The tool stores nothing and takes no action. The call uses the official Anthropic SDK.Validate the result
The tool input is parsed with a strict schema. Unknown fields, bad enumerations, malformed times and out-of-range party sizes are rejected. Refusals and truncated responses count as failures.
Fail closed and quietly
Failures return a code, never raw model output. Request text is not logged or stored.
Keep configuration on the server
The API key and model identifier come from environment variables. Clients can’t choose a model. Requests are same-origin only, size-limited and rate-limited on a best-effort basis.
This boundary is tested with an injected client and sanitised fixtures: valid output, invalid output, wrong tool, truncation, upstream failure, timeout and an injection attempt. The tests make no live API calls. Without credentials, the endpoint reports that it is unavailable; it does not fake a result.
05Why model improvements help
Each improvement maps to a measurable change.
| If Claude gets better at… | The product should show… | Measured as |
|---|---|---|
| Telling firm from flexible | Fewer wrongly dropped candidates and fewer violated limits | Preference satisfaction; hard-constraint violation rate |
| Knowing when to ask | Fewer unnecessary questions without more wrong guesses | Human clarification rate |
| Reliable tool use | Fewer malformed or repeated calls | Tool-call failure recovery |
| Long-context workflows | Coordination that survives cancellations and long gaps | Cancellation recovery rate; time to complete |
| Consistent reasoning | Stable proposals when inputs change slightly | Matching stability |
The deterministic validator caps the cost of a model error. That is what makes it safe to try a new model: re-run the evaluation, compare, adopt. Prompt and model version evaluation is Planned (Intended. Not built.).
06Capabilities in play
What is used, what is being explored.
- Tool use Prototype (Built in the repository. Not used with real players.)
- One forced tool for structured extraction. Multi-tool orchestration is the next question.
- Structured output Prototype (Built in the repository. Not used with real players.)
- Schema-defined tool input, strictly validated before use.
- Long-running workflows Research (A question we are working on. No result is claimed.)
- How to carry state across a multi-day coordination, and what belongs in context versus in the state layer.
- Context handling Research (A question we are working on. No result is claimed.)
- What the model needs to see about a player pool and history, and what it should never see.
- MCP Planned (Intended. Not built.)
- A candidate for exposing scoped tools such as venue inventory. It would be adopted only if it makes permissions and audit simpler, not for its own sake.
- Agent evaluation Research (A question we are working on. No result is claimed.)
- Synthetic scenarios and metric definitions are written. No runs are published.
07Open technical questions
Where deeper expertise would change our decisions.
- Architecture for a production agent that proposes while a deterministic layer disposes.
- Long-running workflows: where state should live, and how to resume safely after a gap.
- Tool use under partial failure: retries, idempotency and recovery.
- Context management for pools of players and prior pairings, with minimal retained free text.
- Structured outputs: how strict to be, and what to do on near-misses.
- Evaluation design: scenario coverage, graders and measuring unnecessary intervention.
- MCP architecture for scoped tools, permissions and auditing.
- Safe action boundaries for irreversible steps.
- Running larger evaluation sweeps across prompt and model versions.
08Summary
In short.
Claude reasons. Deterministic systems decide what is true. Claude interprets intent, proposes groups, selects tools and handles exceptions. Deterministic software owns players, capacity, availability, bookings, payments and the final validation.