Research topic · Multi-agent coordination
Measure the collisions before the pull request
When several coding agents work on one repository, much of the waste happens before any pull request exists: two agents take the same task, one overwrites another, a claim is never released. Pull-request history cannot show this. A coordination log can.
The situation
How should we measure duplicate effort and conflicting actions between agents before the final artefact is reviewed?
What you leave with
Coordination event schema, conflict taxonomy and an inspectable analysis notebook that recomputes every reported rate from the raw events.
For: Agent platform researcher; software engineering lab
What is established, and by whom
Background. Coding agents now produce pull requests at scale, and the coordination paper starts from reports that their pull requests are produced faster but accepted less often. Pull-request telemetry records the end of the process: by then, duplicated effort has already been spent and overwritten work is already gone.
My result. In Before the Pull Request (arXiv preprint), concurrent coding agents coordinated through grite, a git-native append-only coordination log. Three findings:
- With the shared substrate, the share of work that merely re-did a teammate’s task fell from 78% to 0%, while useful throughput more than tripled, at bounded overhead.
- Every agent’s copy of the log converged to the same state with no write silently dropped, whereas a file-based tracker lost concurrent writes.
- The log is minable: conflicting edits, lock starvation, redundant rediscovery and race-to-close were recovered automatically with provenance, several of them invisible in pull-request history.
The disclosure matters here. grite is a Neul Labs project, Neul Labs is a company I founded, and it builds AI agent infrastructure. The paper is therefore evidence about a substrate I have an interest in, produced by me. Treat it as a demonstration of what the measurement can show, not as an independent evaluation of the tool.
What a reproducible coordination study should measure
The rate that matters most is usually not “conflicts” in the abstract but wasted completed work: effort that reached an end state and was discarded because another agent did the same thing. Alongside it:
- Lost writes — events one agent made that no other agent ever observed.
- Claim latency and starvation — how long work waits on a claim that is never released.
- Failure provenance — for every counted failure, the event sequence that produced it, so a reviewer can disagree with a classification.
Illustrative example — not a client engagement. A lab runs four agents on a backlog of 60 issues and suspects overlap. The study fixes definitions first, then compares their current tracker with an append-only log over three repeated runs. The analysis notebook shows most waste comes from two agents re-discovering the same failing test, not from edit conflicts, which points the next experiment at shared test state rather than locking.
Open questions
Whether these rates hold with heterogeneous models, mixed human-agent teams and long-lived repositories is untested. So is the link between less duplication and higher pull-request acceptance: plausible, but a separate measurement. These are the questions I would most like to study with a second, independent substrate.
Useful if
- You run several coding agents on shared work and suspect duplication or overwrites you cannot see in review.
- You are studying agent collaboration and need failure modes that can be counted with provenance, not anecdotes.
- You can instrument the coordination layer, or export its events, for the period you want to study.
Not the right fit if
- You want a coordination layer installed and operated for your team. That is delivery work.
- Your agents never share work, so the question is single-agent quality; see agent failure taxonomies instead.
- You need the study to endorse a particular coordination tool.
What this produces
- Coordination event schema. The events a study needs — claim, release, edit, close, conflict — with actor, timestamp, target and causal link, mapped onto whatever your agents already emit.
- Conflict taxonomy. Agreed definitions for duplicate work, conflicting edits, lock starvation, redundant rediscovery and race-to-close, with the rule that assigns each event sequence to a class.
- Analysis notebook. Recomputes every rate from raw events and links each counted failure back to the events that produced it.
- Findings note. What the rates show, their uncertainty across runs, and which conclusions the design cannot support.
What you need before starting
- Event access. Coordination or tracker logs with timestamps and agent identity, or permission to add an append-only log for the study period.
- A task set. Representative work items, ideally with several agents assigned concurrently, and a fixed agent configuration per arm.
- A baseline. The coordination you use today, so the comparison is against your practice rather than against nothing.
Protocol
- 1 Fix definitions first. Agree what counts as duplicate and conflicting work before any logs are analysed. Some parallel exploration is deliberate and must not be counted as waste.
- 2 Instrument. Capture an append-only event stream per agent; check that every agent's copy converges and that no write is silently dropped.
- 3 Run matched arms. Same tasks, models and agent count, with and without the coordination change, repeated across runs.
- 4 Mine failures. Classify event sequences into the taxonomy and keep provenance for every instance.
- 5 Report with limits. Rates with run-to-run spread, examples of each failure class, and what the setup does not generalise to.
Limits and unfavourable results
- The published result comes from one substrate (grite), one author and a controlled harness. It has not been independently replicated.
- A coordination log measures coordination. It does not tell you whether the merged code is correct or whether reviewers will accept it.
- Conflict definitions involve judgement; changing them changes the rates, which is why they are fixed before analysis.
- grite is a Neul Labs project and I founded Neul Labs. A study can use your own logging instead, and this conflict is disclosed in any write-up.
Evidence behind this page
- Before the Pull Request: Mining Multi-Agent Coordination
2026 · Preprint (not peer reviewed)
- grite (GitHub, Neul Labs) — Git-native, append-only, signed coordination and issue log for agents; the paper's dataset, harness and mining toolkit are released with it. MIT. Neul Labs is a company I founded.
Questions
How do you measure duplicate work between coding agents?
Log every claim, edit and close as an event with agent identity, then count work items that more than one agent carried to completion without building on the other. In Before the Pull Request, a shared append-only log reduced that share from 78% to 0% while useful throughput more than tripled.
Why not just analyse pull requests?
Pull requests show the outcome, not the coordination. Several failure modes in the paper, such as redundant rediscovery and race-to-close, were recoverable from the coordination log but invisible in pull-request history.
Is there a benchmark for multi-agent conflict detection?
I am not aware of an established one. The paper releases a dataset, harness and mining toolkit, which is a starting point; a conflict benchmark would also need agreed definitions and labelled cases from more than one substrate.
Does the study require grite?
No. grite was the instrument in the paper. A study can use your existing tracker or agent logs if they record enough events, or a neutral log added for the study. Because I founded Neul Labs, which builds grite, that choice is discussed at scoping.
See also
Research topic
RAG Knowledge Conflicts and Input Regimes
How conflict detectors for RAG behave across input regimes: five findings from an audit, an input-regime ladder and open questions.
Method
Agent Failure Taxonomies and Incident Coding
Versioned codebook, ambiguity rules, double coding with agreement statistics and annotated examples for agent traces.
Method
Research Reproducibility Packages
Environment manifest, inputs with provenance, per-table scripts, raw per-run outputs and a known-limits statement, with worked examples.
Technical guides
Commission a scoped study
Send the claim, the decision it informs and what access exists. A short, non-confidential description is enough to start; data and code access are agreed afterwards.
Last reviewed 2026-10-07.