NSF grant ~$1.49M: LMGame, an open-source ecosystem that turns games into instruments for AI evaluation (UC San Diego)
An NSF POSE Phase II award of ~$1.49 million to UC San Diego. Games combine clear rules, measurable outcomes, and real-time challenges — instruments for revealing whether AI can reason, plan, perceive, and decide. Yet game-based AI evaluations have been ad hoc and fragmented, with scattered benchmarks and volunteer-maintained testbeds. LMGame unifies diverse game environments in a shared open-source ecosystem with standardized metrics and reproducible results. Runs July 2026 to June 2028.
Grant overview (primary data)
- Award amount$1,494,278
- RecipientUniversity of California-San Diego (California)
- ProgramPOSE
- Period2026-07-15 〜 2028-06-30
- FunderU.S. National Science Foundation (NSF) / NSF
Key points
- Recipient: UC San Diego, ~$1.49M, July 2026 to June 2028 (POSE Phase II)
- Games as measurement instruments for AI: clear rules, measurable outcomes, real-time challenges
- Replaces scattered benchmarks and volunteer-maintained testbeds with a unified LLM-first platform and shared API
- Governance formalized: technical core committee, signed releases, responsible disclosure, reproducible harnesses
- Vision: contamination-aware auditable metrics and reproducible pipelines linking evaluation to post-training
Headlines usually run one way — AI beats a game. This award flips the perspective: games become instruments that measure what AI can and cannot do. Games offer a rare combination for an evaluation device: clear rules, measurable outcomes, and real-time challenges. As chess and Go once served as touchstones for AI, diverse digital games can mirror reasoning, planning, perception, and decision-making.
1Game-based benchmarks were ad hoc
The problem is that this use has been ad hoc. Game-based benchmarks are scattered across repositories, metrics are inconsistent, and testbeds depend on short-term volunteer maintenance by researchers — results are hard to reproduce and harder to compare.
LMGame answers with ecosystem-building: a single LLM-first platform where evaluation and training both run on unified game environments and a shared API; a technical core committee overseeing signed versioned releases and responsible disclosure; reproducible test harnesses for every game and metric; and partner pathways across academia, industry, and non-profits with education and outreach.
The POSE Phase II designation itself is an investment in exactly this transition — from research prototype to sustainably governed open-source project.
2Contamination-aware, auditable metrics
The most forward-looking phrase in the record is auditable, contamination-aware metrics. Benchmark contamination — test items leaking into training data — has steadily eroded trust in static test sets. Games, whose states are generated dynamically from rules, offer structural resistance to that problem.
With reproducible pipelines connecting evaluation to post-training also in scope, the project reads as an attempt to build trustworthy AI evaluation as an open public good.
Why it matters
AI capability evaluation underpins adoption decisions and regulatory compliance, and this is a public investment in the open infrastructure behind it. The structural answer to benchmark contamination and the evaluation-to-post-training pipeline are reference points for practitioners in AI evaluation and model validation.
FAQ
What is the POSE program?
Why are games suited to AI evaluation?
What does contamination-aware mean?
Sources (primary)
Source: NSF Award Search (U.S. National Science Foundation, public domain). Amounts are the obligated amount. For privacy, we do not handle principal investigator names.
- NSF Award (original, official)
- NSF Award ID: 2550249