Yoshua Bengio
Yoshua Bengio (born March 5, 1964) is a Canadian computer scientist whose research has focused on deep learning, representation learning, neural language models, and generative models. He is a full professor in the Department of Computer Science and Operations Research at the Universite de Montreal, the founder and scientific advisor of Mila, and a Canada CIFAR AI Chair.[1][2][6] He is also co-president and scientific director of LawZero, a nonprofit research organization that he founded in 2025.[33][35]
Bengio shared the 2018 ACM A.M. Turing Award with Geoffrey Hinton and Yann LeCun. The Association for Computing Machinery recognized the three "for conceptual and engineering breakthroughs that have made deep neural networks a critical component of computing."[17] Bengio's collaborative research includes early analyses of long-term dependency learning, a neural probabilistic language model, layer-wise training of deep networks, denoising autoencoders, attention-based neural machine translation, generative adversarial networks, and generative flow networks.[8][9][10][11][12][13][16]
Education and academic career
Bengio was born in Paris, France. He earned a Bachelor of Engineering from McGill University in 1986, a master's degree in computer science in 1988, and a PhD in computer science in 1991.[1][4] After his doctorate, he held postdoctoral positions at the Massachusetts Institute of Technology, where he worked on statistical learning and sequential data, and at AT&T Bell Laboratories, where he worked on learning and vision algorithms.[1][3]
He joined the Universite de Montreal faculty in 1993 and is a full professor in its Department of Computer Science and Operations Research.[2][3][4] His work there has included research on recurrent networks, probabilistic models, representation learning, deep architectures, and machine learning safety. Sources that describe his career use several overlapping labels, including artificial intelligence, machine learning, and deep learning. These labels describe related parts of his research record rather than separate appointments.
Mila and Element AI
In 1993, Bengio established a research laboratory at the Universite de Montreal that later developed into Mila. The institute expanded in 2017 through collaboration between the Universite de Montreal and McGill University, with Polytechnique Montreal and HEC Montreal, and became a nonprofit organization in 2018.[5]
Bengio served as Mila's scientific director before moving to a newly created founder and scientific advisor position on March 28, 2025. Mila said the change would allow him to spend more time on AI safety research and international governance work. He remained a core academic member, a member of its Scientific Council, and a Canada CIFAR AI Chair.[6] Institutional profiles that predate this transition may still describe him as Mila's scientific director.
Bengio also co-founded the Montreal company Element AI in 2016.[1][2] ServiceNow completed its acquisition of Element AI on January 8, 2021, according to the company's filing with the United States Securities and Exchange Commission.[7]
Research
Bengio's publications span several periods of neural-network research. Attribution matters because the works below were collaborative and often had a student or colleague as first author.
| Area | Representative work | What the cited work established |
|---|---|---|
| Recurrent networks | "Learning long-term dependencies with gradient descent is difficult" (1994) | Bengio, Patrice Simard, and Paolo Frasconi analyzed why gradient-based learning becomes increasingly difficult as the relevant temporal span grows.[8] |
| Neural language models | "A Neural Probabilistic Language Model" (2003) | Bengio, Rejean Ducharme, Pascal Vincent, and Christian Jauvin jointly learned distributed word representations and a probability model for word sequences.[9] |
| Training deep networks | "Greedy Layer-Wise Training of Deep Networks" (2006) | Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle studied layer-wise pretraining and reported that it improved optimization and generalization in their experiments.[10] |
| Representation learning | "Extracting and Composing Robust Features with Denoising Autoencoders" (2008) | Pascal Vincent, Larochelle, Bengio, and Pierre-Antoine Manzagol trained autoencoders to reconstruct clean inputs from corrupted versions and stacked the learned representations.[11] |
| Neural machine translation | "Neural Machine Translation by Jointly Learning to Align and Translate" (2014) | Dzmitry Bahdanau, Kyunghyun Cho, and Bengio introduced a translation model that learned a soft alignment while generating a target sentence.[12] |
| Generative modeling | "Generative Adversarial Nets" (2014) | Ian Goodfellow and seven coauthors, including Bengio, formulated adversarial training as a minimax game between a generator and a discriminator.[13] |
| Generative flow networks | "GFlowNet Foundations" (2023) | Bengio and five coauthors developed theoretical foundations for models designed to sample diverse structured candidates approximately in proportion to a reward.[16] |
Long-term dependencies and neural language models
The 1994 long-term dependency paper did not introduce a later gated recurrent architecture. Instead, it explained a training difficulty: gradient methods face greater difficulty as the time interval between relevant information and its use increases.[8] The result became part of the technical background for later work on training recurrent neural networks, but later architectures should be credited to their own authors.
The 2003 neural probabilistic language model addressed the curse of dimensionality in language modeling by learning a distributed representation for each word together with a model of word-sequence probability.[9] It is accurate to describe the paper as an influential early neural language model. It is less precise to say that it single-handedly invented word embeddings, since distributed representations and vector-space approaches have a broader history.
Deep architectures and representation learning
In 2006, Bengio and his coauthors examined greedy layer-wise training at a time when deep multilayer networks were difficult to optimize from random initialization. Their experiments compared several pretraining strategies and found that unsupervised layer-wise pretraining could place the network in a region that supported better optimization and generalization.[10] The paper studied both restricted Boltzmann machine and autoencoder variants; it did not claim that layer-wise pretraining was the only way to train deep networks.
The 2008 denoising autoencoder paper trained a model to reconstruct an original input after corruption. The authors used the task to learn features and showed how multiple denoising autoencoders could be composed into a deeper network.[11] In 2015, Bengio, LeCun, and Hinton published a review in Nature that summarized developments in deep learning across supervised, unsupervised, and reinforcement-learning settings.[14] Bengio later coauthored the 2016 textbook Deep Learning with Goodfellow and Aaron Courville.[15]
Attention and generative models
The 2014 translation paper by Bahdanau, Cho, and Bengio replaced a single fixed-length source representation with a mechanism that learned which source positions to emphasize for each generated target word.[12] The paper is a major early result in neural attention for machine translation. It should not be described as the sole origin of all later attention systems, which developed through multiple research lines.
Bengio was one of eight authors of the original 2014 generative adversarial network paper. Goodfellow was the first author. The paper proposed training a generator and discriminator against one another in a two-player minimax game and demonstrated the framework on generated samples.[13] Describing GANs as Bengio's sole invention would therefore misstate the paper's authorship.
In later work, Bengio and collaborators developed generative flow networks, or GFlowNets. The 2023 Journal of Machine Learning Research paper describes them as models that can sample a diverse set of candidates approximately in proportion to a reward and can represent distributions over structured objects such as sets and graphs.[16] The paper presents theoretical properties and training objectives, not a claim that GFlowNets replace every form of probabilistic inference or reinforcement learning.
Turing Award
ACM announced in March 2019 that Bengio, Hinton, and LeCun would share the 2018 Turing Award.[17] The year in the award's name is therefore 2018 even though the public announcement and award lecture took place in 2019.
ACM's account credited Bengio with work on probabilistic sequence models, high-dimensional word representations, neural attention, and generative modeling. It credited Hinton and LeCun with distinct, overlapping lines of neural-network research.[17] The award citation recognizes their combined conceptual and engineering contributions. It does not support referring to any one of the three as the sole inventor of deep learning.
AI safety and governance
In March 2023, Bengio signed a Future of Life Institute open letter calling for a six-month pause in training systems more powerful than GPT-4.[27] In May 2023, he also signed the Center for AI Safety statement that called extinction risk from AI a global priority alongside other societal-scale risks.[28] These records establish Bengio's participation in the two advocacy initiatives. The statements themselves are positions and calls for action, not measurements of the probability of a particular outcome.
Bengio was the lead author of "Managing extreme AI risks amid rapid progress," a policy forum article published in Science in May 2024. The paper discussed large-scale social harms, malicious use, and possible loss of human control. It called for technical safety research and adaptive, proactive governance while also noting a lack of consensus about how some extreme risks would arise or be managed.[29]
He chaired the International Scientific Report on the Safety of Advanced AI, commissioned by the United Kingdom government. The May 2024 interim report was overseen by an expert advisory panel involving 30 countries as well as the European Union and the United Nations.[31] Bengio also chaired the International AI Safety Report 2026, which was written with guidance from more than 100 independent experts, including nominees from more than 30 countries and international organizations.[32] Both reports distinguish evidence, uncertainty, and disagreement rather than assigning Bengio's personal views to every contributor.
In 2026, Bengio and journalist Maria Ressa served as co-chairs of the United Nations Independent International Scientific Panel on Artificial Intelligence. The 40-member panel released a preliminary assessment on July 1, 2026, ahead of the first Global Dialogue on AI Governance.[30][34] The panel's mandate is scientific assessment. Its report should not be described as legislation or as a binding international standard.
LawZero and Scientist AI
Bengio launched LawZero on June 3, 2025, as a nonprofit organization for research on AI systems designed with safety as an explicit objective.[35] As of July 28, 2026, LawZero identifies him as co-president and scientific director, while also listing his roles at the Universite de Montreal and Mila.[33]
The organization's initial technical proposal is called Scientist AI. A February 2025 preprint describes a non-agentic system built around a world model and a question-answering inference machine, both intended to represent uncertainty explicitly.[36] The authors propose that such a system could assist research and provide checks on agentic systems. The paper is a research proposal, not evidence that a production Scientist AI has achieved those goals.
A June 2026 preprint gives formal and semi-formal arguments for a "disinterested" predictor trained to approximate a Bayesian posterior over contextualized statements.[37] Its safety conclusions are conditional on assumptions about the training process, initialization, and the rarity of dangerous predictors. The authors also state that the proposed guarantees do not cover every risk that could arise when the predictor is placed inside a larger agentic system.[37]
Awards and honors
The following list is selective. Dates refer to the conferring body's award year, which can differ from the date of an announcement or ceremony.
| Year | Honor | Scope and attribution |
|---|---|---|
| 2017 | Officer of the Order of Canada | Awarded May 11, 2017, and invested May 10, 2018.[18] |
| 2018 | ACM A.M. Turing Award | Shared with Geoffrey Hinton and Yann LeCun; announced in 2019.[17] |
| 2019 | Killam Prize in Natural Sciences | Awarded to Bengio by the Canadian Killam program.[19] |
| 2020 | Fellow of the Royal Society | The Royal Society records his election in 2020.[20] |
| 2022 | Princess of Asturias Award for Technical and Scientific Research | Shared with Hinton, LeCun, and Demis Hassabis.[21] |
| 2023 | Gerhard Herzberg Canada Gold Medal for Science and Engineering | NSERC records 2023 as the prize year.[22] |
| 2024 | VinFuture Grand Prize | Shared by Bengio, Hinton, LeCun, Jensen Huang, and Fei-Fei Li for contributions associated with deep learning.[23] |
| 2025 | Queen Elizabeth Prize for Engineering | One of seven laureates recognized for contributions to modern machine learning.[24] |
| 2025 | Officer of the National Order of Quebec | Included among the officers admitted in the 2025 appointments.[25] |
| 2025 | International member of the US National Academy of Sciences | Elected as one of the academy's new international members.[26] |
Selected works
| Year | Publication | Authors or editors |
|---|---|---|
| 1994 | "Learning long-term dependencies with gradient descent is difficult" | Yoshua Bengio, Patrice Simard, and Paolo Frasconi.[8] |
| 2003 | "A Neural Probabilistic Language Model" | Yoshua Bengio, Rejean Ducharme, Pascal Vincent, and Christian Jauvin.[9] |
| 2006 | "Greedy Layer-Wise Training of Deep Networks" | Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle.[10] |
| 2008 | "Extracting and Composing Robust Features with Denoising Autoencoders" | Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol.[11] |
| 2014 | "Neural Machine Translation by Jointly Learning to Align and Translate" | Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio.[12] |
| 2014 | "Generative Adversarial Nets" | Ian Goodfellow and seven coauthors, including Yoshua Bengio.[13] |
| 2015 | "Deep learning" | Yann LeCun, Yoshua Bengio, and Geoffrey Hinton.[14] |
| 2016 | Deep Learning | Ian Goodfellow, Yoshua Bengio, and Aaron Courville.[15] |
| 2023 | "GFlowNet Foundations" | Yoshua Bengio, Salem Lahlou, Tristan Deleu, Edward J. Hu, Mo Tiwari, and Emmanuel Bengio.[16] |
| 2024 | "Managing extreme AI risks amid rapid progress" | Yoshua Bengio and more than 20 coauthors.[29] |
See also
- Deep Learning
- Neural Network
- Neural Machine Translation
- Generative adversarial network
- AI safety
- LawZero
- Geoffrey Hinton
- Yann LeCun
- ACM A.M. Turing Award
References
- ^Fundacion Princesa de Asturias. "Geoffrey Hinton, Yann LeCun, Yoshua Bengio and Demis Hassabis: Trajectory." 2022. fpa.es/...n-lecun-yoshua-bengio-and-demis-hassabis
- ^Universite de Montreal, Department of Computer Science and Operations Research. "Yoshua Bengio." diro.umontreal.ca/...in13599
- ^Canadian Artificial Intelligence Association. "Dr. Yoshua Bengio." caiac.ca/...yoshua-bengio
- ^McGill University. "Bengio co-recipient of A.M. Turing Award." March 27, 2019. mcgill.ca/...o-co-recipient-am-turing-award-295735
- ^Mila. "About Mila." mila.quebec/...about-mila
- ^Mila. "Transition in Mila's Scientific Direction: Yoshua Bengio Becomes Scientific Advisor and Laurent Charlin Appointed Interim Scientific Director." March 28, 2025. mila.quebec/...ition-in-milas-scientific-direction
- ^ServiceNow, Inc. "Completion of Acquisition or Disposition of Assets." Form 8-K, filed January 14, 2021. sec.gov/...now-20210108
- ^Yoshua Bengio, Patrice Simard, and Paolo Frasconi. "Learning long-term dependencies with gradient descent is difficult." *IEEE Transactions on Neural Networks* 5, no. 2 (1994): 157-166. pubmed.ncbi.nlm.nih.gov/18267787
- ^Yoshua Bengio, Rejean Ducharme, Pascal Vincent, and Christian Jauvin. "A Neural Probabilistic Language Model." *Journal of Machine Learning Research* 3 (2003): 1137-1155. jmlr.org/...bengio03a
- ^Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle. "Greedy Layer-Wise Training of Deep Networks." *Advances in Neural Information Processing Systems* 19 (2006). proceedings.neurips.cc/...aeb2fae32403405-Abstract
- ^Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol. "Extracting and Composing Robust Features with Denoising Autoencoders." *Proceedings of the 25th International Conference on Machine Learning* (2008): 1096-1103. cs.toronto.edu/...-2008-denoising-autoencoders.pdf
- ^Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. "Neural Machine Translation by Jointly Learning to Align and Translate." arXiv:1409.0473, 2014; presented at ICLR 2015. arxiv.org/...1409.0473
- ^Ian J. Goodfellow et al. "Generative Adversarial Nets." *Advances in Neural Information Processing Systems* 27 (2014). proceedings.neurips.cc/...9a61f95710dbe25-Abstract
- ^Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. "Deep learning." *Nature* 521 (2015): 436-444. nature.com/...nature14539
- ^Ian Goodfellow, Yoshua Bengio, and Aaron Courville. *Deep Learning*. MIT Press, 2016. deeplearningbook.org
- ^Yoshua Bengio, Salem Lahlou, Tristan Deleu, Edward J. Hu, Mo Tiwari, and Emmanuel Bengio. "GFlowNet Foundations." *Journal of Machine Learning Research* 24, no. 210 (2023): 1-55. jmlr.org/...22-0364
- ^Association for Computing Machinery. "Fathers of the Deep Learning Revolution Receive ACM A.M. Turing Award." March 27, 2019. awards.acm.org/...turing-award-2018.pdf
- ^Governor General of Canada. "Mr. Yoshua Bengio: Officer of the Order of Canada." gg.ca/...146-3280
- ^Universite de Montreal Faculty of Arts and Science. "Killam Prize." fas.umontreal.ca/...prix-killam
- ^Royal Society. "Professor Yoshua Bengio OC FRS." royalsociety.org/...yoshua-bengio-25320
- ^Fundacion Princesa de Asturias. "2022 Princess of Asturias Award for Technical & Scientific Research." 2022. fpa.es/...n-lecun-yoshua-bengio-and-demis-hassabis
- ^Natural Sciences and Engineering Research Council of Canada. "Yoshua Bengio." nserc-crsng.canada.ca/...yoshua-bengio
- ^VinFuture Foundation. "The 2024 VinFuture Prize honors four scientific works under the theme of 'Resilient Rebound'." December 6, 2024. vinfutureprize.org/...e-theme-of-resilient-rebound
- ^Queen Elizabeth Prize for Engineering. "2025 QEPrize Winners: Modern Machine Learning." qeprize.org/...modern-machine-learning
- ^Government of Quebec. "Ordre national du Quebec, 2025 insignia ceremony." June 18, 2025. quebec.ca/...c-accueille-de-nouveaux-membres-63777
- ^Universite de Montreal. "Yoshua Bengio elected international member of the National Academy of Sciences." May 1, 2025. nouvelles.umontreal.ca/...onal-academy-of-sciences
- ^Yoshua Bengio. "Slowing down development of AI systems passing the Turing test." April 5, 2023. yoshuabengio.org/...ai-systems-passing-turing-test
- ^Center for AI Safety. "Statement on AI Risk." 2023. safe.ai/...statement-on-ai-risk
- ^Yoshua Bengio et al. "Managing extreme AI risks amid rapid progress." *Science* 384, no. 6698 (2024): 842-845. arxiv.org/...2310.17688
- ^Independent International Scientific Panel on Artificial Intelligence. *Preliminary Report*. United Nations, July 2026. un.org/...en_Preliminary%20Report_.pdf
- ^United Kingdom Department for Science, Innovation and Technology and AI Safety Institute. "International Scientific Report on the Safety of Advanced AI: interim report." May 17, 2024. gov.uk/...ific-report-on-the-safety-of-advanced-ai
- ^International AI Safety Report. "2026 Report: Executive Summary." February 3, 2026. internationalaisafetyreport.org/...ecutive-summary
- ^LawZero. "Yoshua Bengio." lawzero.org/...yoshua-bengio
- ^United Nations. "Media advisory: Preliminary Report of the Independent International Scientific Panel on Artificial Intelligence." July 1, 2026. un.org/...c%20Panel%20Report%20July%201-Latest.pdf
- ^LawZero. "Yoshua Bengio Launches LawZero: A New Nonprofit Advancing Safe-by-Design AI." June 3, 2025. lawzero.org/...-nonprofit-advancing-safe-design-ai
- ^Yoshua Bengio et al. "Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?" arXiv:2502.15657, February 2025. arxiv.org/...2502.15657
- ^Yoshua Bengio et al. "Safety from Honesty in a Disinterested AI Predictor." arXiv:2606.29657, June 2026. arxiv.org/...2606.29657
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
8 revisions · v9 · 2,903 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent 2026-07-28 fact-check: 37 explicit scholarly, institutional, regulatory, government, award-body, and United Nations references; 75 resolved citation calls; 10 canonical internal targets; and 23 high-risk root source groups checked. Root inspected all 36 desktop/mobile captures plus selected source-evidence renders. Verified identity, education, current roles, Mila and Element AI chronology, paper-specific attribution, award dates, international-report and UN-panel roles, LawZero boundaries, and conditional Scientist AI preprint claims. Protected-shorter preservation review passed for the exact candidate.
Cite this page: AI Wiki. "Yoshua Bengio." aiwiki.ai, updated 30 Jul 2026, fact-checked 30 Jul 2026. CC BY 4.0. https://aiwiki.ai/wiki/yoshua_bengio