Research

Can an agent turn fragmented intent into a completed group, without breaking a hard rule?

Doubles runs a small, applied coordination research effort alongside the product. It is product research, not academic research: the aim is to know what is safe to automate.

Research (A question we are working on. No result is claimed.) Everything below is a question, a definition or a design. No results, scores or benchmarks have been produced.

The question, and the ones underneath it.

Can frontier agents reliably turn fragmented human intent into completed real-world group actions while respecting hard constraints?

  • When should model reasoning hand off to deterministic optimization?
  • How should hard and soft human preferences be represented?
  • How should the system recover when one participant cancels?
  • How do we measure coordination quality?
  • How much clarification is necessary before an agent can act?
  • When should an agent ask a question rather than infer a preference?
  • How can a system avoid repeatedly grouping the same people?
  • How should uncertain skill levels affect matching?
  • How do we measure unnecessary intervention?

Propose, then validate.

Research (A question we are working on. No result is claimed.) A system in which a model proposes and a deterministic validator disposes can make hard-constraint violations at execution impossible by construction. If so, the remaining research is about proposal quality: completion, clarification and intervention. Our working belief, to be tested, is that most of the value is there rather than in validation.

This is a hypothesis. It could be wrong, for instance if validation turns out to reject so many proposals that completion suffers.

Nine evaluation dimensions.

Definitions, not results
DimensionThe question it answers
Hard-constraint violation rateHow often does a proposed or executed action break a rule that cannot be broken?
Group completion rateOf the groups the system sets out to form, how many reach four confirmed players?
Human clarification rateHow often must a player be asked a question before the system can act?
Cancellation recovery rateWhen a confirmed player drops, how often is the group restored without a human?
Manual intervention rateHow often does a human operator have to step in?
Matching stabilityDoes a small change in input cause a large, unjustified change in the proposed group?
Preference satisfactionHow many stated soft preferences are honoured in the final group?
Time to complete coordinationHow long from first request to a confirmed game?
Tool-call failure recoveryWhen a tool call fails or times out, does the workflow recover safely?

Hard-constraint violation rate is the gating metric: no autonomy is extended while it is above zero on the test set. Everything else is traded against it, never the reverse.

Scenario families.

Research (A question we are working on. No result is claimed.) Scenarios are written by hand and generated variants, each with a known correct outcome, so grading does not depend on opinion.

Contradictory preferences
A request whose own conditions cannot all hold.
Ambiguous time and place
Daypart words, “around” a city, “the other side”.
Uneven skill pool
Pools where the obvious group has a wide skill spread.
Incomplete groups
Three compatible players, or five for four seats.
Cancellation cascades
A drop, then a failed replacement, then another drop.
Stale inventory
A court that changes state between proposal and action.
Repeated pairing
Pools that tempt the system to group the same people again.
Tool failures
Timeouts, partial success and duplicate calls.
Hostile text
Requests that contain instructions aimed at the system.

How real coordination feeds the evaluation.

Padel games in Delhi NCR are arranged by hand today. That work is the source of real phrasing and real failure cases. The intended loop, which is Planned (Intended. Not built.), is:

  1. Collect, with consent

    Keep structured intents and outcomes from real coordination. Do not keep more free text than the evaluation needs.

  2. Replay

    Run those cases through the agent offline and grade against what actually happened.

  3. Shadow

    Let the agent propose alongside the human organiser without acting. Compare proposals.

  4. Extend autonomy narrowly

    Only for action classes where the hard-violation rate has been shown to be zero on the test set.

Seven ways coordination fails.

A. Interpretation
The model misreads a time, area, skill band or flexibility.
B. Constraint
An action would break a hard rule. Must be caught by the validator; counts as a violation if it is not.
C. Coordination
A valid group that is a poor one: unstable, repetitive or needlessly unmatched.
D. Process
Asks too much, asks too little, or involves a human unnecessarily.
E. Tooling
Timeouts, partial writes, duplicate actions, stale reads.
F. Adversarial and privacy
Injected instructions, over-retention of text, data shown to the wrong player.
G. Human factors
No-shows and late changes that no system can prevent, only absorb.

What we do not know.

  • How many requests a human organiser can take before shadow-mode comparison becomes statistically meaningful.
  • Whether players prefer being asked one question or being offered two options.
  • How much a model’s uncertainty about skill should widen the allowed spread.
  • Whether the architecture transfers to activities with different group sizes. Not tested.

The implemented pieces so far are the intent boundary and the validator; see Technology and Evidence.