GlucoFM
GlucoFM is a self-supervised foundation model for continuous glucose monitoring (CGM) data developed by Google Research. It learns representations of 24-hour glucose traces by separating slower trends from shorter deviations and training on both masked context and temporal change. The researchers first posted the preprint on May 29, 2026, updated it on August 25, and published a Google Research summary on August 26.[1][2]
GlucoFM is a research prototype for retrospective physiological representation learning. It is not a diagnostic test, treatment system, or cleared medical device. The paper evaluates whether its frozen representations support study-defined classification and forecasting tasks; it does not report a prospective clinical deployment.[1]
CGM sensors sample glucose in interstitial fluid beneath the skin and provide an estimate associated with blood glucose, usually as a sequence of readings and trends.[3] GlucoFM is intended to encode those sequences for later machine-learning tasks. It does not itself collect sensor data or prescribe an action.[1]
Key facts
| Field | Detail |
|---|---|
| Developer | Google Research[1][2] |
| First preprint | May 29, 2026[1] |
| Current paper discussed here | arXiv:2605.30865v2, August 25, 2026[1] |
| Model type | Self-supervised, dual-stream CGM foundation model[1] |
| Input unit | 24-hour CGM window aligned to a 5-minute grid with 288 positions[1] |
| Trainable parameters | 0.72 million[1] |
| Total parameters during pretraining | 1.18 million, including the exponential-moving-average target branch[1] |
| Pretraining corpus | 109,066 hours from 477 dataset-defined subjects or recording entries[1] |
| Main downstream study | Four cohorts, 203 dataset-defined participants or entries, 71,669 CGM hours, seven unique phenotype tasks, and 14 cohort-task evaluations[1] |
| Reported average PR-AUC | 58.8, compared with 54.7 for the strongest controlled CGM-specific comparator in the study[1][2] |
| Evidence status | Retrospective research reported by the model's authors; no prospective clinical validation or regulatory clearance[1] |
Background
CGM produces dense time-series data that can include fasting patterns, overnight variation, meal-related excursions, missing intervals, and device artifacts. Clinical labels such as laboratory measurements or diagnoses are much less frequent and costlier to collect. This imbalance has encouraged self-supervised models that first learn from unlabeled CGM sequences and are then evaluated on smaller labeled cohorts.[1]
GlucoFM followed several CGM-specific pretrained models. CGMformer used masked learning on daily glucose profiles.[5] GluFormer used autoregressive next-token prediction on more than 10 million measurements from 10,812 mostly non-diabetic adults.[4] The GlucoFM authors also evaluated general time-series models and study-retrained versions of CGM-JEPA, X-CGM-JEPA, and GluFormer.[1]
The distinction between published models and the study's implementations matters. For CGM-JEPA, X-CGM-JEPA, and GluFormer, the GlucoFM team reported that official pretrained checkpoints were unavailable for its controlled comparison. The team reproduced the published methods and retrained them on the same unlabeled corpus used for GlucoFM. Its numerical comparisons therefore describe the authors' implementations under one protocol, not a universal ranking of all versions of those models.[1]
Architecture
Chronological grid and missingness
GlucoFM divides a continuous recording segment into 24-hour windows and aligns each window to a five-minute chronological grid. The resulting input has 288 positions. An observation mask records which values were physically measured. Missing positions can be filled for tensor construction, but the mask remains available to the model so an imputed position is not treated as an observed reading.[1]
The preprocessing treats a gap longer than one hour as a break between recording segments. Shorter gaps remain inside a segment as missing observations. During pretraining, the authors sampled overlapping 24-hour windows to cover different daily start times. Downstream evaluation instead used non-overlapping windows.[1]
State and event streams
The model applies a learnable causal Gaussian filter to estimate a slow-varying state component. It defines an event component from shorter deviations around that local trend. The streams are encoded separately and then fused. "State" and "event" are model representations rather than direct measurements of a person's physiological state, meal timing, activity, or sensor error.[1]
This split is the main difference between GlucoFM and the single-stream CGM models used as comparators. The paper's ablation study reported task-averaged PR-AUC values of 55.3 for a raw-input encoder, 55.6 for state only, 51.8 for event only, and 58.8 for the dual-stream design. Those results support the architecture within the paper's experiments, but they do not establish that the decomposition is optimal for every CGM dataset.[1]
Pretraining objectives
GlucoFM uses two latent-prediction objectives related to Joint Embedding Predictive Architecture methods.[1]
| Objective | Role in the paper |
|---|---|
| Masked contextual representation learning | Masks temporal patches in the online branch and predicts their target-branch latent representations from the surrounding daily context[1] |
| Temporal dynamics modeling | Predicts how the state and event tokens change from one temporal patch to the next[1] |
The model also uses CGM-specific augmentations. Value perturbations simulate effects such as baseline drift and short compression-like drops. Structural sparsification simulates lower sampling density and brief disconnections. In a paper ablation, the complete augmentation setup produced higher task-averaged scores than no augmentation. Preserving the observation mask also performed better than the tested dense-interpolation variants.[1]
The online and target encoders each contain three Transformer layers with a hidden dimension of 128, four attention heads, and a feed-forward dimension of 256. An exponential-moving-average target branch supplies latent targets during pretraining. This branch accounts for much of the difference between 0.72 million trainable parameters and 1.18 million total pretraining parameters. The downstream experiments discard the target branch and freeze the online encoder.[1]
Training data
The authors report 109,066 hours of unlabeled CGM recordings from 477 dataset-defined subjects or recording entries across five sources.[1]
| Pretraining source | Dataset-defined subjects or entries | CGM hours | Sampling interval |
|---|---|---|---|
| Wear-CGM | 192 | 75,330 | 5 minutes[1] |
| ShanghaiT2DM subset | 44 | 12,414 | 15 minutes[1] |
| Stanford subset | 19 | 8,761 | 5 minutes[1] |
| BIG IDEAs | 16 | 3,017 | 5 minutes[1] |
| Colas | 206 | 9,544 | 5 minutes[1] |
Wear-CGM, the largest part of the corpus, is non-public. The paper describes it as two Google and Fitbit research phases involving healthy, non-diabetic adults in the United States. It reports institutional review board approval and written consent for de-identified secondary research and algorithm development.[1]
The count of 477 should not be interpreted as 477 verified unique biological participants. ShanghaiT2DM uses recording visits as subject entries because some people had more than one session. The authors' identity audit found one biological participant with an earlier unlabeled visit in the pretraining subset and a later labeled visit in the downstream subset. For ShanghaiT2DM, separation was therefore defined at the recording-visit level.[1]
The paper reports pretraining GlucoFM for 120 epochs on one NVIDIA H100 GPU with a batch size of 128. That is the experiment's training configuration, not a published hardware requirement for all later use.[1]
Downstream evaluation
Cohorts and labels
The principal evaluation used four cohorts with 203 dataset-defined participants or recording entries and 71,669 hours of CGM data.[1]
| Cohort | Entries | Monitoring data | Study-defined tasks |
|---|---|---|---|
| CGMacros | 45 | 10,376 Dexcom hours and 10,998 Libre hours | Diabetes risk, insulin resistance, obesity, hyperlipidemia[1] |
| ShanghaiT2DM | 65 sessions from 58 biological participants | 15,634 hours | Hypoglycemia, insulin resistance, hyperlipidemia[1] |
| Stanford | 37 | 27,571 hours | Diabetes risk, beta-cell dysfunction, insulin resistance[1] |
| Hall | 56 | 7,090 hours | Diabetes risk, glucotype, insulin resistance, hyperlipidemia[1] |
Together, the cohorts supplied seven unique phenotype-classification tasks and 14 cohort-task combinations. Labels came from cohort metadata, clinical records, or thresholds chosen in the paper. The authors state that those thresholds were used to define consistent research labels and were not intended as standalone diagnostic criteria.[1]
Linear-probe results
For the main comparison, the researchers froze each encoder, extracted representations for non-overlapping 24-hour windows, and trained the same logistic-regression classifier. They used ten iterations of five-fold subject-grouped cross-validation, keeping all windows for one study subject or recording-entry unit in the same fold.[1]
Across the 14 evaluations, GlucoFM's task-averaged precision-recall area under the curve (PR-AUC) was 58.8. CGM-JEPA, the strongest controlled CGM-specific comparator on that aggregate, scored 54.7. The paper reports a paired average difference of 4.11 points with a 95% confidence interval from 2.40 to 5.81 points.[1] Google summarized the same result as a 4.1-point absolute gain, or about 7.5% relative to that baseline.[2]
The authors reported that GlucoFM had the highest PR-AUC in every diabetes-risk and beta-cell-dysfunction evaluation and in three of four insulin-resistance evaluations. Its advantages were less consistent for other tasks, and the paper identified ShanghaiT2DM as a difficult setting because of a device and sampling-rate shift.[1]
These figures measure retrospective discrimination of study-defined labels. They do not show that GlucoFM can diagnose a condition, improve a treatment decision, or produce calibrated risk estimates in a clinical population.[1][7]
Postprandial response forecasting
A separate experiment examined two-hour glucose changes after logged meals. It used 874 paired meal events from 34 CGMacros participants, with Dexcom and Libre modeled separately under the same subject splits. The frozen 24-hour pre-meal representation was combined step by step with one hour of recent CGM, meal nutrition, fasting glucose, BMI, and diabetes status.[1]
With the complete context, GlucoFM had a two-sensor mean trajectory error of 21.88 mg/dL, compared with 22.90 mg/dL for CGM-JEPA. The paper describes this as a 4.5% reduction. It also reports the lowest full-context errors among the evaluated methods for positive incremental area under the curve, peak rise, and peak time.[1][2]
The full-context qualification is important. The frozen GlucoFM representation by itself did not lead every postprandial endpoint. Its strongest result appeared after the study added recent CGM, meal information, and participant variables.[1]
Transfer, limited labels, and multiple days
The paper trained linear probes on one cohort and tested them directly on another for shared diabetes-risk and insulin-resistance tasks. GlucoFM ranked first on 21 of 24 reported PR-AUC and ROC-AUC transfer evaluations. It did not lead every endpoint, including some transfers from Stanford to CGMacros for insulin resistance.[1]
In the authors' few-shot experiments, GlucoFM had the highest task-averaged PR-AUC at each tested number of support subjects and each tested observation fraction. In a separate multiday analysis, averaging daily embeddings improved most settings as more days were added. ShanghaiT2DM insulin resistance was an exception under simple mean pooling, which the authors used to show that aggregation can depend on the cohort and label.[1]
Limitations and clinical status
The v2 paper identifies the pretraining population as modest and potentially unrepresentative of wider demographic, disease, device, and lifestyle variation. Its largest data source is non-public. The encoder still processes each day independently, so its multiday experiments rely on simple downstream aggregation rather than native modeling across weeks or months.[1]
All reported evaluations were retrospective. The study did not test prospective clinical outcomes, treatment response, medication or insulin dosing, real-time operation, or integration into a care workflow. It also did not report an independent reproduction by a separate research group. The paper says the authors plan to release code and reproducibility scripts; it does not present that planned release as completed.[1]
The authors state that GlucoFM has not been cleared or approved by a regulatory authority and is not intended to diagnose, treat, cure, or prevent disease or to replace professional medical advice.[1] This is consistent with the distinction between a research representation model and deployed AI in healthcare.
The American Diabetes Association's 2026 standards recommend CGM in several diabetes-management settings and stress user and caregiver education.[6] The same standards state that evidence is insufficient to use CGM for screening or diagnosis of prediabetes or diabetes.[7] GlucoFM's retrospective classification results do not change those clinical recommendations and should not be interpreted as medical advice.
References
- ^Zechen Li et al. "GlucoFM: A Dual-Stream Foundation Model for Continuous Glucose Monitoring." arXiv:2605.30865v2, Aug. 25, 2026. arxiv.org/...2605.30865
- ^Ahmed A. Metwally and Zechen Li. "GlucoFM: Foundation model for continuous glucose monitoring." Google Research, Aug. 26, 2026. research.google/...r-continuous-glucose-monitoring
- ^U.S. Food and Drug Administration. "What is the pancreas? What is an artificial pancreas device system?" Accessed Aug. 27, 2026. fda.gov/...-what-artificial-pancreas-device-system
- ^Guy Lutsker et al. "A foundation model for continuous glucose monitoring data." Nature 650, 978-986, Jan. 14, 2026. doi.org/...s41586-025-09925-9
- ^Yang Lu et al. "A pretrained transformer model for decoding individual glucose dynamics from continuous glucose monitoring data." National Science Review 12(5), nwaf039, 2025. doi.org/...nwaf039
- ^American Diabetes Association Professional Practice Committee. "7. Diabetes Technology: Standards of Care in Diabetes-2026." Diabetes Care 49, Supplement 1, 2026. doi.org/...dc26-s007
- ^American Diabetes Association Professional Practice Committee. "2. Diagnosis and Classification of Diabetes: Standards of Care in Diabetes-2026." Diabetes Care 49, Supplement 1, 2026. doi.org/...dc26-s002
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 2,141 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independently fact-checked against the cited sources on Aug. 27, 2026; claims were limited to what those sources support.
Cite this page: AI Wiki. "GlucoFM." aiwiki.ai, updated 27 Aug 2026, fact-checked 27 Aug 2026. CC BY 4.0. https://aiwiki.ai/wiki/glucofm