$1,494,278 POSE

NSF grant ~$1.49M: LMGame, an open-source ecosystem that turns games into instruments for AI evaluation (UC San Diego)

University of California-San Diego California Started Jul 2026

An NSF POSE Phase II award of ~$1.49 million to UC San Diego. Games combine clear rules, measurable outcomes, and real-time challenges — instruments for revealing whether AI can reason, plan, perceive, and decide. Yet game-based AI evaluations have been ad hoc and fragmented, with scattered benchmarks and volunteer-maintained testbeds. LMGame unifies diverse game environments in a shared open-source ecosystem with standardized metrics and reproducible results. Runs July 2026 to June 2028.

Grant overview (primary data)

  • Award amount$1,494,278
  • RecipientUniversity of California-San Diego (California)
  • ProgramPOSE
  • Period2026-07-15 〜 2028-06-30
  • FunderU.S. National Science Foundation (NSF) / NSF

Key points

  • Recipient: UC San Diego, ~$1.49M, July 2026 to June 2028 (POSE Phase II)
  • Games as measurement instruments for AI: clear rules, measurable outcomes, real-time challenges
  • Replaces scattered benchmarks and volunteer-maintained testbeds with a unified LLM-first platform and shared API
  • Governance formalized: technical core committee, signed releases, responsible disclosure, reproducible harnesses
  • Vision: contamination-aware auditable metrics and reproducible pipelines linking evaluation to post-training

Headlines usually run one way — AI beats a game. This award flips the perspective: games become instruments that measure what AI can and cannot do. Games offer a rare combination for an evaluation device: clear rules, measurable outcomes, and real-time challenges. As chess and Go once served as touchstones for AI, diverse digital games can mirror reasoning, planning, perception, and decision-making.

1Game-based benchmarks were ad hoc

The problem is that this use has been ad hoc. Game-based benchmarks are scattered across repositories, metrics are inconsistent, and testbeds depend on short-term volunteer maintenance by researchers — results are hard to reproduce and harder to compare.

LMGame answers with ecosystem-building: a single LLM-first platform where evaluation and training both run on unified game environments and a shared API; a technical core committee overseeing signed versioned releases and responsible disclosure; reproducible test harnesses for every game and metric; and partner pathways across academia, industry, and non-profits with education and outreach.

The POSE Phase II designation itself is an investment in exactly this transition — from research prototype to sustainably governed open-source project.

2Contamination-aware, auditable metrics

The most forward-looking phrase in the record is auditable, contamination-aware metrics. Benchmark contamination — test items leaking into training data — has steadily eroded trust in static test sets. Games, whose states are generated dynamically from rules, offer structural resistance to that problem.

With reproducible pipelines connecting evaluation to post-training also in scope, the project reads as an attempt to build trustworthy AI evaluation as an open public good.

Why it matters

AI capability evaluation underpins adoption decisions and regulatory compliance, and this is a public investment in the open infrastructure behind it. The structural answer to benchmark contamination and the evaluation-to-post-training pipeline are reference points for practitioners in AI evaluation and model validation.

FAQ

What is the POSE program?
Pathways to Enable Open-Source Ecosystems — an NSF program that helps research software grow into sustainably governed open-source ecosystems. Phase II supports the transition to a mature project with formal governance and development practices.
Why are games suited to AI evaluation?
Rules are explicit, outcomes are measurable, and play demands real-time decisions — so reasoning, planning, and perception can be scored. Because game states are generated dynamically, evaluations can also resist training-data contamination.
What does contamination-aware mean?
When benchmark items leak into training data, a model can appear capable while merely reciting memorized answers. Contamination-aware metrics are designed to detect and account for that possibility while measuring ability.

Sources (primary)

Source: NSF Award Search (U.S. National Science Foundation, public domain). Amounts are the obligated amount. For privacy, we do not handle principal investigator names.

#AI#NSF#Research grant#Benchmark#Open source#LLM#UC San Diego
Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.