Gemma 4 Developer Agent Competition
The Gemma 4 Developer Agent Competition is a 2026 Kaggle challenge for AI coding agents, accompanied by a research-paper track. Kaggle lists Google DeepMind as host, while the rules name Google LLC as sponsor.[2][3] Google Research announced a combined $100,000 prize pool on October 2, 2026, describing the aim as adapting open models for software-development work on everyday hardware.[4]
Tracks and prizes
Paper-track entry does not require entering the prediction competition.[3]
| Track | Award | Prize |
|---|---|---|
| Prediction competition | First place | $37,000 |
| Prediction competition | Second place | $18,000 |
| Prediction competition | Third place | $10,000 |
| Paper track | Overall Best Paper | $15,000 |
| Paper track | Best New Resource | $10,000 |
| Paper track | Best New Application | $10,000 |
The rules allocate $65,000 to prediction prizes and $35,000 to paper awards.[2][5]
Schedule
The published schedule distinguishes entry from final submission.[1][3]
| Event | Date in 2026 |
|---|---|
| Paper track opens | September 22 |
| Prediction competition opens | September 23 |
| Paper submission deadline | November 12 |
| Prediction entry and team-merger deadline | November 25 |
| Final prediction submission deadline | December 2 |
Deadlines are 23:59 UTC unless otherwise specified; organizers may revise them.[1][3]
Tasks and dataset
The public development set contains 129 Python bug-fix and feature-request tasks from FastAPI, Requests, Rich and HTTPX. Each task supplies an issue description, a repository commit, a reference patch and verification tests. Repository snapshots exclude subsequent commit history.[6]
| Resource | Contents |
|---|---|
tasks.jsonl | Issue descriptions, identifiers, reference fixes and tests |
snapshots/ | Git working trees frozen before the fixes |
graphs/ | Directed code graphs representing symbols and their relationships |
embeddings/ | 256-dimensional vectors for semantic code retrieval |
wheels/ | Python dependencies for installation without internet access |
The graph resources support dependency exploration and retrieval of related code. During scoring, hidden tasks replace the development set. Organizers describe approximately 120 test tasks from private repositories, divided equally between public and private leaderboard splits.[6]
The curation process checks that a test reproduces a failure before the reference fix and that the full suite passes after it. Test-set curation also checks whether a larger frontier model can pass, or come within one failing test of passing, the case.[6]
Agent package and evaluation
An entry is a submission.zip archive containing a root agent.yaml. The harness converts this configuration into an agent.[1] The underlying Google Agent Development Kit uses YAML Agent Config files to describe instructions, tools and subagents.[9]
Every agent and subagent must use gemma-4-31b-it-qat-w4a16-ct. LoRA adapters are optional; different subagents may select different adapters.[1] Google's linked model card identifies the checkpoint as a Gemma 4 quantization-aware-training release in compressed-tensors format for vLLM, distributed under Apache 2.0.[7]
Only harness tools or custom agent_tool subagents are allowed; paths cannot escape the submission root. The shared runtime allowance is 12 hours, including sandbox setup but excluding patch validation.[1]
Each issue receives a pass/fail result. The overall score is the percentage of patched repositories passing their validation tests.[1]
The available L4x4 notebook option provides 96 GB of GPU memory. These sessions require internet access to be disabled and consume quota at twice the T4x2/P100 rate.[1]
Paper requirements and judging
The paper track accepts original, unpublished research of at most 3,000 words. Entries use a Kaggle Writeup covering the abstract, introduction, methods, experiments, related work and citations; code is optional. A public paper PDF can replace full text in the Writeup. Unsubmitted drafts are ineligible.[3]
Judges score five equally weighted criteria from 0 to 5 and average the results:
| Criterion | Focus |
|---|---|
| Novelty | New insights or understanding |
| Quality | Generalization beyond the competition |
| Relevance | Contribution to software engineering and agentic learning |
| Verifiability | Enough detail to understand methods and evidence |
| Clarity | Clear presentation |
The track is non-archival, so entering does not prevent submission to another conference or venue.[3]
Participation and release requirements
Both tracks permit teams of up to five people and up to two final submissions for judging. The prediction competition allows one submission per day; the paper track allows five.[2][5]
The rules identify Apache 2.0 as the winner-license type. Winners must provide reproducible code and documentation, including training and inference procedures and the required computing environment. External data and tools are subject to accessibility, licensing and cost conditions, rather than unrestricted use.[2]
Relationship to SWE-bench
SWE-bench, introduced by Carlos E. Jimenez and colleagues, evaluates patches against tests for repository-level issues. Its original release contained 2,294 problems from 12 Python repositories; resolving them can require edits across functions or files.[8] The competition's separately curated development and hidden test sets are distinct from the original SWE-bench dataset.[6][8]
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Google DeepMind. "Google - The Gemma 4 Developer Agent Competition: overview, evaluation, timeline and harness rules." Kaggle. Accessed October 4, 2026.
- ^1 ^2 ^3 ^4Google LLC. "Gemma 4 Developer Agent Competition: rules." Kaggle. Accessed October 4, 2026.
- ^1 ^2 ^3 ^4 ^5 ^6Google DeepMind. "Google - The Gemma 4 Developer Agent Paper Track: overview." Kaggle. Accessed October 4, 2026.
- ^Google Research. "Gemma 4 Developer Agent Competition announcement." X, October 2, 2026. Official embedded-post text retrieved October 4, 2026.
- ^1 ^2Google LLC. "Gemma 4 Developer Agent Paper Track: rules." Kaggle. Accessed October 4, 2026.
- ^1 ^2 ^3 ^4Google DeepMind. "Gemma 4 Developer Agent Competition: dataset description." Kaggle. Accessed October 4, 2026.
- ^Google. "Gemma 4: gemma-4-31b-it-qat-w4a16-ct." Kaggle model card. Accessed October 4, 2026.
- ^1 ^2Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press and Karthik Narasimhan. "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" ICLR 2024; arXiv:2310.06770, revised November 11, 2024.
- ^Google. "Build agents with Agent Config." Agent Development Kit documentation. Accessed October 4, 2026.
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 982 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent source review on October 4, 2026. Full new article independently reviewed against rendered official overview, data description, separate track rules, model card and academic SWE-bench context. Entry, paper and final-submission dates distinguished; no hidden harness or future outcome claims.
Cite this page: AI Wiki. "Gemma 4 Developer Agent Competition." aiwiki.ai, updated 3 Oct 2026, fact-checked 3 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/gemma_4_developer_agent_competition