What happens when people are paired with AI — looked at individually rather than on average, results fell into three modes
The effect of human-AI collaboration is usually reported as a single average. Looking at individuals, this preliminary study finds results falling into three modes, and finds that what separated them was neither model benchmarks nor raw cognitive ability.
Paper overview (our summary)
- Field (arXiv category)cs.CY(+1)
- AuthorsVivienne Ming
- Submitted2026-07-02
- arXiv ID2607.02467v1
Key points
- A preliminary brief examining human-AI collaboration at the level of the individual rather than as an average.
- Results fell into three modes: deferring to the model, leaning on it to endorse an answer already settled on, and genuine complementary reasoning.
- Those who leaned on the model to endorse an answer already settled on performed worse than the model alone.
- What separated the modes was not raw cognitive ability or model benchmarks but perspective-taking, intellectual humility and curiosity.
- A real-money prediction market served as an externally resolved benchmark, and a pre-registered replication is in preparation.
1What an average hides
Whether pairing people with AI raises or lowers performance is often reported as one average. An average, though, returns a single number even when quite different outcomes are mixed within it. This study opens that mixture by looking at individuals.
- 1The first modeDeferring to the model, ending level with it
- 2The second modeLeaning on the model to endorse an answer already settled on, ending worse than the model alone
- 3The third modeGenuine complementary reasoning, matching or exceeding the market itself
- 4The shapeThree modes, not representable by a single average
The second mode is the one to note. The model was used, and the result was worse than not using it. Where a model is treated as material to confirm what one already thought, an error in judgment is reinforced rather than corrected. Holding a tool and learning from it are different things.
2What separated the modes
On the author account, what distinguished who reached the third mode was not intelligence or model performance but traits bearing on collaboration: whether one can imagine a view unlike one own, whether one can hold that one might be wrong, whether one wants to check. These are also qualities that training and design can work on.
3An externally resolved benchmark
Using a market where money moves and outcomes later resolve is itself notable: the researcher does not decide what counts as correct. That said, this is a preliminary brief and the author states a replication is being prepared. It is better read as a way of framing the question than as a settled conclusion.
4The link to other research on this site
This site separately covers work sorting out how many kinds of human-AI team are even being studied. The point that averages cannot carry the story, and the point that one term covers different things, both warn against speaking of these matters in aggregate. This article is our own summary of public research information and does not warrant its contents.
Why it matters
Reporting the effect of human-AI collaboration as one average conceals the group for whom it backfires. If the separating factor is collaborative disposition, there is room to work on how a tool is used rather than only on adopting it.
FAQ
Can using AI make performance worse?
How settled is this result?
Sources (primary)
Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.
- arXiv abstract page (original, official)
- PDF (arXiv)
- arXiv ID: 2607.02467