Citation and evidence

Gemma 4 Developer Agent Competition

5 min full readUpdated 9 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI Code GenerationAI ResearchGoogle DeepMind

Cite this article

The Gemma 4 Developer Agent Competition is a 2026 Kaggle challenge for AI coding agents, accompanied by a research-paper track. Kaggle lists Google DeepMind as host, while the rules name Google LLC as sponsor.[2][3] Google Research announced a combined $100,000 prize pool on October 2, 2026, describing the aim as adapting open models for software-development work on everyday hardware.[4]

Tracks and prizes

Paper-track entry does not require entering the prediction competition.[3]

TrackAwardPrize
Prediction competitionFirst place$37,000
Prediction competitionSecond place$18,000
Prediction competitionThird place$10,000
Paper trackOverall Best Paper$15,000
Paper trackBest New Resource$10,000
Paper trackBest New Application$10,000

Expanded article table

The rules allocate $65,000 to prediction prizes and $35,000 to paper awards.[2][5]

Schedule

The published schedule distinguishes entry from final submission.[1][3]

EventDate in 2026
Paper track opensSeptember 22
Prediction competition opensSeptember 23
Paper submission deadlineNovember 12
Prediction entry and team-merger deadlineNovember 25
Final prediction submission deadlineDecember 2

Expanded article table

Deadlines are 23:59 UTC unless otherwise specified; organizers may revise them.[1][3]

Tasks and dataset

The public development set contains 129 Python bug-fix and feature-request tasks from FastAPI, Requests, Rich and HTTPX. Each task supplies an issue description, a repository commit, a reference patch and verification tests. Repository snapshots exclude subsequent commit history.[6]

ResourceContents
tasks.jsonlIssue descriptions, identifiers, reference fixes and tests
snapshots/Git working trees frozen before the fixes
graphs/Directed code graphs representing symbols and their relationships
embeddings/256-dimensional vectors for semantic code retrieval
wheels/Python dependencies for installation without internet access

Expanded article table

The graph resources support dependency exploration and retrieval of related code. During scoring, hidden tasks replace the development set. Organizers describe approximately 120 test tasks from private repositories, divided equally between public and private leaderboard splits.[6]

The curation process checks that a test reproduces a failure before the reference fix and that the full suite passes after it. Test-set curation also checks whether a larger frontier model can pass, or come within one failing test of passing, the case.[6]

Agent package and evaluation

An entry is a submission.zip archive containing a root agent.yaml. The harness converts this configuration into an agent.[1] The underlying Google Agent Development Kit uses YAML Agent Config files to describe instructions, tools and subagents.[9]

Every agent and subagent must use gemma-4-31b-it-qat-w4a16-ct. LoRA adapters are optional; different subagents may select different adapters.[1] Google's linked model card identifies the checkpoint as a Gemma 4 quantization-aware-training release in compressed-tensors format for vLLM, distributed under Apache 2.0.[7]

Only harness tools or custom agent_tool subagents are allowed; paths cannot escape the submission root. The shared runtime allowance is 12 hours, including sandbox setup but excluding patch validation.[1]

Each issue receives a pass/fail result. The overall score is the percentage of patched repositories passing their validation tests.[1]

The available L4x4 notebook option provides 96 GB of GPU memory. These sessions require internet access to be disabled and consume quota at twice the T4x2/P100 rate.[1]

Paper requirements and judging

The paper track accepts original, unpublished research of at most 3,000 words. Entries use a Kaggle Writeup covering the abstract, introduction, methods, experiments, related work and citations; code is optional. A public paper PDF can replace full text in the Writeup. Unsubmitted drafts are ineligible.[3]

Judges score five equally weighted criteria from 0 to 5 and average the results:

CriterionFocus
NoveltyNew insights or understanding
QualityGeneralization beyond the competition
RelevanceContribution to software engineering and agentic learning
VerifiabilityEnough detail to understand methods and evidence
ClarityClear presentation

Expanded article table

The track is non-archival, so entering does not prevent submission to another conference or venue.[3]

Participation and release requirements

Both tracks permit teams of up to five people and up to two final submissions for judging. The prediction competition allows one submission per day; the paper track allows five.[2][5]

The rules identify Apache 2.0 as the winner-license type. Winners must provide reproducible code and documentation, including training and inference procedures and the required computing environment. External data and tools are subject to accessibility, licensing and cost conditions, rather than unrestricted use.[2]

Relationship to SWE-bench

SWE-bench, introduced by Carlos E. Jimenez and colleagues, evaluates patches against tests for repository-level issues. Its original release contained 2,294 problems from 12 Python repositories; resolving them can require edits across functions or files.[8] The competition's separately curated development and hidden test sets are distinct from the original SWE-bench dataset.[6][8]

References

  1. ^1 ^2 ^3 ^4 ^5 ^6 ^7Google DeepMind. "Google - The Gemma 4 Developer Agent Competition: overview, evaluation, timeline and harness rules." Kaggle. Accessed October 4, 2026.
  2. ^1 ^2 ^3 ^4Google LLC. "Gemma 4 Developer Agent Competition: rules." Kaggle. Accessed October 4, 2026.
  3. ^1 ^2 ^3 ^4 ^5 ^6Google DeepMind. "Google - The Gemma 4 Developer Agent Paper Track: overview." Kaggle. Accessed October 4, 2026.
  4. ^Google Research. "Gemma 4 Developer Agent Competition announcement." X, October 2, 2026. Official embedded-post text retrieved October 4, 2026.
  5. ^1 ^2Google LLC. "Gemma 4 Developer Agent Paper Track: rules." Kaggle. Accessed October 4, 2026.
  6. ^1 ^2 ^3 ^4Google DeepMind. "Gemma 4 Developer Agent Competition: dataset description." Kaggle. Accessed October 4, 2026.
  7. ^Google. "Gemma 4: gemma-4-31b-it-qat-w4a16-ct." Kaggle model card. Accessed October 4, 2026.
  8. ^1 ^2Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press and Karthik Narasimhan. "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" ICLR 2024; arXiv:2310.06770, revised November 11, 2024.
  9. ^Google. "Build agents with Agent Config." Agent Development Kit documentation. Accessed October 4, 2026.

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

v1 · 982 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independent source review on October 4, 2026. Full new article independently reviewed against rendered official overview, data description, separate track rules, model card and academic SWE-bench context. Entry, paper and final-submission dates distinguished; no hidden harness or future outcome claims.

Cite this page: AI Wiki. "Gemma 4 Developer Agent Competition." aiwiki.ai, updated 3 Oct 2026, fact-checked 3 Oct 2026. CC BY 4.0. https://aiwiki.ai/wiki/gemma_4_developer_agent_competition

Suggest edit

What links here