Acronyms

RawGraph

Acronyms is a practical glossary of abbreviations, shortened names, and recurring initialisms used across artificial intelligence, machine learning, statistics, robotics, data systems, and adjacent fields. In strict usage, an acronym is formed from initial letters and pronounced as a word, while an initialism is spoken letter by letter.[1][2] AI writing does not consistently observe that distinction, so this page follows common technical usage and includes both.

The same letters can mean different things across subfields. For example, BN can mean batch normalization or Bayesian network, SSL can mean self-supervised learning or semi-supervised learning, and SVD usually means singular value decomposition but can mean singing voice detection in music information retrieval. Read each entry in the context of its paper, software package, dataset, or application domain. Definitions below are compact identifiers, not substitutes for the linked topic articles or original sources.

See also: Guides, Terms, and Abbreviations.

How to use this glossary

  • Capitalization and punctuation can be meaningful. CoT, GloVe, Grad-CAM, i.i.d., and reCAPTCHA preserve source or conventional styling.
  • A slash or semicolon marks a genuine context-dependent expansion, not interchangeable wording inside one method.
  • Product and project names such as CatBoost, Gurobi, and SWI-Prolog appear because readers encounter them as shortened technical names, even when they are not strict acronyms.
  • The glossary favors stable, source-backed expansions. It merges duplicate rows and removes malformed or unsupported expansions from the earlier version.

A

TermMeaning
A*A* Search Algorithm
A3CAsynchronous Advantage Actor-Critic
ABACAttribute-Based Access Control [4]
A/B TestingA statistical method for comparing two or more treatments or algorithms [3]
ACEAlternating conditional expectation algorithm
ACOAnt Colony Optimization
AdamOptimization algorithm named for adaptive moment estimation; Adam is not a strict initialism [32]
ADASYNAdaptive Synthetic Sampling [75]
ADTAutomatic Drum Transcription
AEAutoencoder
AGCAdaptive Gradient Clipping [76]
AGIArtificial general intelligence
AIArtificial intelligence
AIaaSArtificial Intelligence as a Service
ALActive Learning
AMActivation maximization [77]
AMRAbstract Meaning Representation
AMTAutomatic music transcription [72]
ANIArtificial Narrow Intelligence
ANNArtificial neural network
ANOVAAnalysis of variance
APIApplication Programming Interface
ARAugmented reality
ASIArtificial superintelligence
ASICApplication-Specific Integrated Circuit
ASRAutomatic speech recognition
ASTAutomated speech translation
AUCArea under the curve, usually the ROC curve unless another curve is named
AutoMLAutomated Machine Learning [3]

B

TermMeaning
BB84A quantum key distribution protocol (named after its inventors, Bennett and Brassard, and the year 1984) [59]
BBOBiogeography-Based Optimization
BCEBinary cross-entropy
BDTBoosted Decision Tree
BERTBidirectional Encoder Representations from Transformers [10]
BFSBreadth-First Search
BIBusiness Intelligence
BiFPNBidirectional Feature Pyramid Network [134]
BILSTMBidirectional Long Short-Term Memory
BLEUBilingual evaluation understudy [46]
BNBatch normalization; also Bayesian network
BNNBayesian neural network; also binarized neural network
BOBayesian Optimization
BPBackpropagation
BPEByte Pair Encoding [67]
BPMFBayesian Probabilistic Matrix Factorization [78]
BPNBackpropagation Neural Network
BPTTBackpropagation through time
BQMLBigQuery ML [74]
BRBest-Response (in game theory)
BRDFBidirectional reflectance distribution function
BRNNBidirectional Recurrent Neural Network
BRRBayesian ridge regression

C

TermMeaning
CADComputer-Aided Design
CAEContractive Autoencoder
CAMClass activation mapping in computer vision; also computer-aided manufacturing
CAPTCHACompletely Automated Public Turing test to tell Computers and Humans Apart [79]
CARTClassification and Regression Trees
CASEComputer-Aided Software Engineering
CatBoostCategorical Boosting [80]
CAVConcept activation vector
CBACContent-Based Access Control
CBOWContinuous Bag of Words
CBRCase-Based Reasoning
CCACanonical Correlation Analysis
CCCConcordance correlation coefficient; canonical correlation coefficient in some multivariate contexts
CCECategorical cross-entropy
CECross-Entropy
CECConstant Error Carousel
CEGARCounterexample-Guided Abstraction Refinement
CEGISCounterexample-Guided Inductive Synthesis
CFCollaborative filtering
cGANConditional Generative Adversarial Network [137]
CLConfident learning; also contrastive learning and continual learning [98][99][51]
CLIPContrastive Language-Image Pre-Training [14]
CLNNConditionaL neural network, using the source paper's stylized capitalization [60]
CMACovariance Matrix Adaptation
CMACCerebellar Model Articulation Controller
CMA-ESCovariance Matrix Adaptation Evolution Strategy
CNNConvolutional neural network
COIN-ORComputational Infrastructure for Operations Research
ConvNetConvolutional Neural Network
CoTChain-of-thought [28]
COTECollective of Transformation-Based Ensembles [81]
CoT promptingChain-of-thought prompting [28]
CPConstraint Programming
CPLEXIBM mathematical optimization solver [130]
CPNColored Petri Nets
CRBMConditional Restricted Boltzmann Machine
CRFConditional Random Field
CRNNConvolutional Recurrent Neural Network
CSLRContinuous Sign Language Recognition
CSPConstraint Satisfaction Problem
CSVComma-separated values
CTCConnectionist Temporal Classification
CT-LSTMContinuous-time long short-term memory; the abbreviation is paper-dependent
CTRCollaborative Topic Regression [82]
CUDACompute Unified Device Architecture [7]
CVComputer vision; cross-validation; or coefficient of variation, depending on context
CycCycL and OpenCyc, a knowledge representation and reasoning system [83]

D

TermMeaning
D*Dynamic A* Search Algorithm
DaaSData as a Service
DAEDenoising AutoEncoder or Deep AutoEncoder
DAMLDARPA Agent Markup Language
DARTDropouts meet Multiple Additive Regression Trees, a tree-boosting regularization method [31]
DBMDeep Boltzmann Machine
DBNDeep belief network
DBSCANDensity-Based Spatial Clustering of Applications with Noise [84]
DCAIData-centric AI
DCGANDeep Convolutional Generative Adversarial Network [138]
DDPGDeep Deterministic Policy Gradient
DEDifferential evolution
DeconvNetDeConvolutional Neural Network
DeepLIFTDeep Learning Important FeaTures [25]
DFSDepth-First Search
DLDeep learning
DMData mining; also diffusion model in some machine learning literature
DNNDeep neural network
DPDynamic Programming
DPODirect preference optimization [37]
DQNDeep Q-network [34]
DRDetection Rate
DRLDeep Reinforcement Learning
DSData Science
DSRDeep symbolic regression [85]
DSRLDeep symbolic reinforcement learning [61]
DSSDecision Support System
DTDecision Tree
DTDDeep Taylor Decomposition [86]
DWTDiscrete Wavelet Transform

E

TermMeaning
EDAExploratory data analysis
EKFExtended Kalman Filter
ELECTRAEfficiently Learning an Encoder that Classifies Token Replacements Accurately [13]
ELMExtreme Learning Machine [87]
ELMoEmbeddings from Language Models [11]
ELUExponential Linear Unit [88]
EMExpectation maximization
EMDEarth mover's distance; also empirical mode decomposition and entropy-minimization discretization
ERNIEEnhanced Representation through kNowledge IntEgration [89]
ESEvolution Strategies
ESNEcho State Network
ETLExtract, Transform, Load
EXTExtra-Trees, or extremely randomized trees

F

TermMeaning
F1F1 score, the harmonic mean of precision and recall
FALAFinite Action-set Learning Automata
Fast R-CNNFast Region-based Convolutional Network [20]
FCFully-Connected
FCMFuzzy C-Means
FCNFully Convolutional Network
FERFacial Expression Recognition
FFTFast Fourier transform
FLFederated Learning
FLOPFloating-point operation
FLOPSFloating-point operations per second, more clearly written FLOP/s
FMFoundation model
FNFalse negative
FNNFeedforward Neural Network
FNRFalse negative rate
FOAFFriend of a Friend (ontology)
FPFalse positive
FPGAField-Programmable Gate Array
FPNFeature Pyramid Network [21]
FPRFalse positive rate
FSLFew-shot learning
FSTFinite state transducer
FTLFederated transfer learning [90]
FWAFireworks Algorithm
FWIoUFrequency Weighted Intersection over Union

G

TermMeaning
GAGenetic Algorithm
GALEGlobal Aggregations of Local Explanations [91]
GAMGeneralized Additive Model
GANGenerative Adversarial Network [17]
GAPGlobal Average Pooling
GBDTGradient Boosted Decision Tree
GBMGradient Boosting Machine
GCNGraph convolutional network [52]
GDGradient descent
GEBIGlobal Explanation for Bias Identification [92]
GLCMGray Level Co-occurrence Matrix
GLMGeneralized Linear Model
GLOMA neural network architecture by Geoffrey Hinton [93]
Gloss2TextA task of transforming raw glosses into meaningful sentences.
GloVeGlobal Vectors for Word Representation [65]
GLPKGNU Linear Programming Kit
GLUEGeneral Language Understanding Evaluation [50]
GMMGaussian mixture model
GNNGraph neural network
GPGaussian process; also genetic programming
GPRGaussian process regression
GPTGenerative Pre-trained Transformer [9]
GPUGraphics processing unit
GQAGrouped-Query Attention [45]
Grad-CAMGradient-weighted class activation mapping [27]
GRUGated recurrent unit [94]
GurobiAn optimization solver (named after its founders, Zonghao Gu, Edward Rothberg, and Robert Bixby)

H

TermMeaning
HamNoSysHamburg Sign Language Notation System [73]
HANHierarchical Attention Networks
HCHierarchical Clustering
HDPHierarchical Dirichlet process [95]
HFHugging Face [96]
hLDAHierarchical Latent Dirichlet allocation
HMMHidden Markov Model
HNNHopfield Neural Network
HOGHistogram of Oriented Gradients (feature descriptor)
HPCHigh Performance Computing
HREDHierarchical Recurrent Encoder-Decoder
HRIHuman-Robot Interaction
HSMMHidden Semi-Markov Model

I

TermMeaning
IaaSInfrastructure as a Service [5]
ICAIndependent component analysis
ICPIterative Closest Point (point cloud registration)
ID3Iterative Dichotomiser 3
IDA*Iterative Deepening A* Search Algorithm
IGIntegrated gradients [97]
i.i.d.Independently and identically distributed
IIDIndependently and identically distributed
ILASPInductive Learning of Answer Set Programs [100]
ILPInteger linear programming; also inductive logic programming
INFDExplanation Infidelity [101]
IoAInternet of Agents [102]
IoEInternet of Everything
IoTInternet of Things
IoUJaccard index (intersection over union)
IRInformation Retrieval
IRCoTInterleaving Retrieval CoT [131]
ISICInternational Skin Imaging Collaboration
IVRInteractive Voice Response

J

TermMeaning
JEPAJoint-embedding predictive architecture [41]

K

TermMeaning
KANKolmogorov-Arnold network [40]
KBKnowledge Base
KDEKernel Density Estimation
KFKalman Filter
kFCVK-fold cross validation
KLKullback-Leibler divergence
K-MeansK-Means Clustering [3]
KNNK-nearest neighbors
KRKnowledge Representation
KRRKernel Ridge Regression

L

TermMeaning
LAIONLarge-scale Artificial Intelligence Open Network [103]
LAMALAnguage Model Analysis
LaMDALanguage Models for Dialog Applications [15]
LBPLocal Binary Pattern (texture descriptor)
LDALatent Dirichlet allocation; also linear discriminant analysis [66][135]
LEPORLength Penalty, Precision, n-gram Position difference Penalty and Recall [49]
LightGBMLight Gradient Boosting Machine [104]
LIMELocal Interpretable Model-agnostic Explanations [105]
LINGOA software for linear, nonlinear, and integer optimization
LLLifelong learning
LLMLarge language model [3]
LLSLinear least squares
LMNNLarge Margin Nearest Neighbor [106]
LoRALow-rank adaptation [35]
LPLinear Programming
LRPLayer-wise Relevance Propagation
LSALatent semantic analysis
LSILatent Semantic Indexing
LSTMLong short-term memory
LSTM-CRFLong Short-Term Memory with Conditional Random Field
LTRLearning To Rank
LVQLearning Vector Quantization

M

TermMeaning
M2MMachine to Machine
MADEMasked Autoencoder for Distribution Estimation [107]
MAEMean absolute error
MAFMasked Autoregressive Flows [108]
MAIRLMulti-Agent Inverse Reinforcement Learning
MAPMaximum A Posteriori (MAP) Estimation
MAPEMean absolute percentage error
MARLMulti-Agent Reinforcement Learning
MARTMultiple Additive Regression Trees [31]
MaxEntMaximum Entropy
MAXSATMaximum Satisfiability Problem
MCLNNMasked ConditionaL Neural Networks [60]
MCMCMarkov Chain Monte Carlo
MCPModel Context Protocol [54]
MCTSMonte Carlo Tree Search
MDLMinimum description length (MDL) principle
MDNMixture Density Network
MDPMarkov Decision Process
MDRNNMultidimensional recurrent neural network
MERMusic Emotion Recognition
METEORMetric for Evaluation of Translation with Explicit ORdering [48]
MHAMulti-head attention [68]
MILMultiple Instance Learning
MILPMixed-Integer Linear Programming
MIoUMean Intersection over Union
MIPMixed-Integer Programming
MLMachine learning
MLAMulti-head latent attention [44]
MLaaSMachine Learning as a Service
MLEMaximum Likelihood Estimation
MLLMMultimodal large language model
MLMMasked language modeling or masked language model
MLPMulti-Layer Perceptron
MMIMaximum Mutual Information
MNISTModified National Institute of Standards and Technology database [3]
MoAMixture of Agents [109]
MoEMixture of Experts [3]
MOEAMulti-Objective Evolutionary Algorithm
MPAMean Pixel Accuracy
MQAMulti-Query Attention [45]
MRMixed Reality
MRFMarkov Random Field
MRRMean Reciprocal Rank
MRSMusic Recommender System
MSEMean squared error
MSRMusic Style Recognition
MTLMulti-Task Learning

N

TermMeaning
NARXNonlinear AutoRegressive with eXogenous input (neural network model)
NASNeural Architecture Search [3]
NBNaive Bayes
NDCGNormalized Discounted Cumulative Gain
NENash Equilibrium (in game theory)
NEATNeuroEvolution of Augmenting Topologies [110]
NERNamed entity recognition
NESTNeural Simulation Tool [111]
NFNormalizing Flow
NFLNo Free Lunch (NFL) theorem
NISQNoisy intermediate-scale quantum
NLGNatural Language Generation
NLPNatural Language Processing
NLUNatural Language Understanding
NMFNon-negative matrix factorization
NMSNon Maximum Suppression
NMTNeural Machine Translation
NNNeural network
NRMSENormalized RMSE
NSGA-IINon-dominated Sorting Genetic Algorithm II [112]
NSTNeural style transfer
NTMNeural Turing Machine [113]
NuSVCNu-Support Vector Classification
NuSVRNu-Support Vector Regression

O

TermMeaning
OCROptical character recognition
ODObject Detection
ODFOnset Detection Function
OILOntology Inference Layer
OLROrdinary Linear Regression
OLSOrdinary Least Squares
OMNeT++Objective Modular Network Testbed in C++
OMROptical music recognition
OOFOut-of-fold
ORBOriented FAST and Rotated BRIEF (feature descriptor)
OWLWeb Ontology Language [6]

P

TermMeaning
PAPixel Accuracy
PaaSPlatform as a Service [5]
PaLMPathways Language Model [16]
PBACPolicy-Based Access Control
PCAPrincipal component analysis
PCLPoint Cloud Library (3D perception)
PEFTParameter-efficient fine-tuning [70]
PEGASUSPre-training with Extracted Gap-sentences for Abstractive Summarization [114]
PFParticle Filter
PLSIProbabilistic Latent Semantic Indexing
PMProject Manager
PMFProbabilistic Matrix Factorization
PMIPointwise Mutual Information
PNNProbabilistic Neural Network
POCProof of Concept
POMDPPartially Observable Markov Decision Process
POSPart of Speech (POS) Tagging
PPLPerplexity (a measure of language model performance)
PPMIPositive Pointwise Mutual Information
PPOProximal Policy Optimization [33]
PReLUParametric rectified linear unit [115]
PRMProbabilistic Roadmap (motion planning algorithm)
PSOParticle Swarm Optimization
PUPositive-unlabeled learning [116]
PYTMPitman-Yor topic model

Q

TermMeaning
QAQuestion Answering
QAOAQuantum Approximate Optimization Algorithm
QAPQuadratic Assignment Problem
QECQuantum Error Correction
QFTQuantum Fourier Transform
QIPQuantum Information Processing
QKDQuantum Key Distribution
QLoRAQuantized low-rank adaptation [36]
QMLQuantum Machine Learning
QNNQuantum Neural Network
QPQuadratic Programming
QPEQuantum Phase Estimation

R

TermMeaning
R2R-squared
RAGRetrieval-Augmented Generation [30]
RandNNRandom Neural Network
RANSACRANdom SAmple Consensus
RBACRole-based access control
RBFRadial Basis Function
RBFNNRadial Basis Function Neural Network
RBMRestricted Boltzmann Machine
R-CNNRegion-based Convolutional Neural Network [20]
RDFResource Description Framework [6]
ReActReasoning and acting [29]
REALMRetrieval-Augmented Language Model Pre-Training [132]
reCAPTCHAGoogle's reCAPTCHA challenge service; "reverse CAPTCHA" is not its expansion [58]
ReLURectified Linear Unit [3]
REPTreeReduced Error Pruning Tree
RETRORetrieval-Enhanced Transformer [133]
RFRandom forest
RFERecursive Feature Elimination
RGBRed Green Blue color model
RICNNRotation Invariant Convolutional Neural Network
RIMRecurrent inference machine [62]
RIPPERRepeated Incremental Pruning to Produce Error Reduction
RISERandom Interval Spectral Ensemble; also Randomized Input Sampling for Explanation [117]
RLReinforcement learning
RLAIFReinforcement learning from AI feedback [38]
RLHFReinforcement Learning from Human Feedback [69]
RMSERoot mean squared error
RMSLERoot mean squared logarithmic error
RMSpropRoot Mean Square Propagation
RNNRecurrent neural network
RNNLMRecurrent Neural Network Language Model (RNNLM)
RoBERTaRobustly Optimized BERT Pretraining Approach [118]
ROCReceiver operating characteristic
ROIRegion Of Interest
RoPERotary position embedding [39]
ROSRobot Operating System [119]
ROUGERecall-Oriented Understudy for Gisting Evaluation (NLP metric) [47]
RPARobotic Process Automation
RRRidge Regression
RRTRapidly-exploring Random Tree (motion planning algorithm)
RSIRecursive self-improvement [120]
RTRLReal-Time Recurrent Learning

S

TermMeaning
SASimulated annealing
SaaSSoftware as a Service [5]
SACSoft Actor-Critic
SAESparse autoencoder; also stacked autoencoder
SAMSegment Anything Model, introduced in Segment Anything [24]
SARSAState-Action-Reward-State-Action
SATBoolean satisfiability problem
SBACSituation-Based Access Control
SBMStochastic block model
SBOStructured Bayesian optimization
SBSESearch-based software engineering
SCIPSolving Constraint Integer Programs
SDAEStacked denoising autoencoder [140]
seq2seqSequence to Sequence Learning
SERSentence Error Rate
SFTSupervised fine-tuning [71]
SGBoostStochastic Gradient Boosting
SGDStochastic gradient descent [3]
SGVBStochastic Gradient Variational Bayes
SHAPSHapley Additive exPlanations [26]
SIFTScale-Invariant Feature Transform (feature detection)
SLSupervised learning
SLAMSimultaneous Localization and Mapping
SLDSSwitching Linear Dynamical System
SLMSmall Language Model
SLPSingle-Layer Perceptron
SLTSign Language Translation [73]
SMA*Simplified Memory-bounded A* Search Algorithm
SMBOSequential Model-Based Optimization
SMOSequential Minimal Optimization
SMOTESynthetic Minority Over-sampling Technique [121]
SNNSpiking neural network; also sparse neural network
SOMSelf-Organizing Map
SOTAState of the Art
SPARQLSPARQL Protocol and RDF Query Language [6]
SPMSentencePiece Model (subword tokenization) [63]
SpRAySpectral Relevance Analysis [122]
SSDSingle Shot MultiBox Detector
SSLSelf-supervised learning; also semi-supervised learning
SSMState space model [42]
STStyle transfer
STaRSelf-Taught Reasoner [123]
STDPSpike Timing-Dependent Plasticity
STLSelf-taught learning [124]
SUMOSimulation of Urban MObility [55]
SURFSpeeded-Up Robust Features (feature detection)
SVCSupport Vector Classification
SVDSingular value decomposition; also singing voice detection
SVMSupport vector machine
SVRSupport Vector Regression
SVSSinging Voice Separation
SWI-PrologAn implementation of the Prolog programming language [57]

T

TermMeaning
T5Text-To-Text Transfer Transformer [12]
TDTemporal Difference
TDATopological data analysis [125]
TDETemporal Dictionary Ensemble [139]
tf-idfterm frequency-inverse document frequency
THAIDTHeta Automatic Interaction Detection
TLTransfer Learning
TNTrue negative
TNRTrue negative rate
ToMTheory of Mind
ToTTree of thoughts [53]
TPTrue positive
TPOTTree-based Pipeline Optimization Tool
TPRTrue positive rate
TPUTensor Processing Unit [8]
TRPOTrust Region Policy Optimization
TSTime series; also tabu search
TSFTime Series Forest [81]
t-SNEt-distributed stochastic neighbor embedding [126]
TSPTraveling Salesman Problem
TTSText-to-Speech

U

TermMeaning
UCTUpper Confidence bounds applied to Trees (Monte Carlo Tree Search variant)
UDAUnsupervised Data Augmentation [127]
UKFUnscented Kalman Filter
ULUnsupervised learning
ULMFiTUniversal Language Model Fine-Tuning
UMAPUniform Manifold Approximation and Projection [64]
USMUniversal Speech Model

V

TermMeaning
VADVoice Activity Detection
VAEVariational AutoEncoder [18]
VGGVisual Geometry Group
VHREDVariational Hierarchical Recurrent Encoder-Decoder [128]
VISSIMVerkehr In Stadten - SIMulationsmodell, a traffic microsimulation system [56]
ViTVision Transformer [23]
VLAVision-language-action model [43]
VLMVision-Language Model
V-NetA fully convolutional network for volumetric medical image segmentation [141]
VQEVariational Quantum Eigensolver
VQ-VAEVector-quantized variational autoencoder [136]
VRVirtual reality
VRPVehicle Routing Problem
VUIVoice User Interface

W

TermMeaning
WCSPWeighted Constraint Satisfaction Problem
WERWord Error Rate
WFSTWeighted finite-state transducer (WFST)
WGANWasserstein Generative Adversarial Network [19]
WMAWeighted Majority Algorithm
WPEWeighted Prediction Error

X

TermMeaning
XAIExplainable Artificial Intelligence
XGBoosteXtreme Gradient Boosting [129]
XORExclusive OR (a common problem in neural networks)

Y

TermMeaning
YOLOYou Only Look Once [22]

Z

TermMeaning
ZSLZero-Shot Learning

References

  1. ^Cambridge Dictionary. "acronym." dictionary.cambridge.org/...acronym
  2. ^Cambridge Dictionary. "initialism." dictionary.cambridge.org/...initialism
  3. ^Google for Developers. "Machine Learning Glossary." developers.google.com/...glossary
  4. ^NIST. "Attribute Based Access Control." csrc.nist.gov/...attribute_based_access_control
  5. ^Mell, P., and Grance, T. (2011). "The NIST Definition of Cloud Computing." csrc.nist.gov/...final
  6. ^World Wide Web Consortium. "Semantic Web Standards." w3.org/...semanticweb
  7. ^NVIDIA. "CUDA Documentation." docs.nvidia.com/cuda
  8. ^Google Cloud. "Cloud TPU documentation." cloud.google.com/...docs
  9. ^Eloundou, T., Manning, S., Mishkin, P., and Rock, D. (2023). "GPTs are GPTs: An early look at the labor market impact potential of large language models." openai.com/...gpts-are-gpts
  10. ^Devlin, J., et al. (2019). "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding." aclanthology.org/N19-1423
  11. ^Peters, M. E., et al. (2018). "Deep contextualized word representations." aclanthology.org/N18-1202
  12. ^Raffel, C., et al. (2020). "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer." jmlr.org/...20-074
  13. ^Clark, K., et al. (2020). "ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators." arxiv.org/...2003.10555
  14. ^Radford, A., et al. (2021). "Learning Transferable Visual Models From Natural Language Supervision." openai.com/...clip
  15. ^Thoppilan, R., et al. (2022). "LaMDA: Language Models for Dialog Applications." arxiv.org/...2201.08239
  16. ^Chowdhery, A., et al. (2022). "PaLM: Scaling Language Modeling with Pathways." arxiv.org/...2204.02311
  17. ^Goodfellow, I., et al. (2014). "Generative Adversarial Nets." papers.nips.cc/...5423-generative-adversarial-nets
  18. ^Kingma, D. P., and Welling, M. (2013). "Auto-Encoding Variational Bayes." arxiv.org/...1312.6114
  19. ^Arjovsky, M., Chintala, S., and Bottou, L. (2017). "Wasserstein Generative Adversarial Networks." proceedings.mlr.press/...arjovsky17a
  20. ^Girshick, R. (2015). "Fast R-CNN." openaccess.thecvf.com/...ast_R-CNN_ICCV_2015_paper
  21. ^Lin, T.-Y., et al. (2017). "Feature Pyramid Networks for Object Detection." openaccess.thecvf.com/..._Networks_CVPR_2017_paper
  22. ^Redmon, J., et al. (2016). "You Only Look Once: Unified, Real-Time Object Detection." openaccess.thecvf.com/...Only_Look_CVPR_2016_paper
  23. ^Dosovitskiy, A., et al. (2020). "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale." arxiv.org/...2010.11929
  24. ^Kirillov, A., et al. (2023). "Segment Anything." arxiv.org/...2304.02643
  25. ^Shrikumar, A., Greenside, P., and Kundaje, A. (2017). "Learning Important Features Through Propagating Activation Differences." proceedings.mlr.press/...shrikumar17a
  26. ^Lundberg, S. M., and Lee, S.-I. (2017). "A Unified Approach to Interpreting Model Predictions." papers.nips.cc/...o-interpreting-model-predictions
  27. ^Selvaraju, R. R., et al. (2017). "Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization." openaccess.thecvf.com/...lanations_ICCV_2017_paper
  28. ^Wei, J., et al. (2022). "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models." arxiv.org/...2201.11903
  29. ^Yao, S., et al. (2022). "ReAct: Synergizing Reasoning and Acting in Language Models." arxiv.org/...2210.03629
  30. ^Lewis, P., et al. (2020). "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." papers.neurips.cc/...780e1bc26945df7481e5-Abstract
  31. ^Rashmi, K. V., and Gilad-Bachrach, R. (2015). "DART: Dropouts meet Multiple Additive Regression Trees." proceedings.mlr.press/...korlakaivinayak15
  32. ^Kingma, D. P., and Ba, J. (2014). "Adam: A Method for Stochastic Optimization." arxiv.org/...1412.6980
  33. ^Schulman, J., et al. (2017). "Proximal Policy Optimization Algorithms." arxiv.org/...1707.06347
  34. ^Mnih, V., et al. (2015). "Human-level control through deep reinforcement learning." nature.com/...nature14236
  35. ^Hu, E. J., et al. (2021). "LoRA: Low-Rank Adaptation of Large Language Models." arxiv.org/...2106.09685
  36. ^Dettmers, T., et al. (2023). "QLoRA: Efficient Finetuning of Quantized LLMs." arxiv.org/...2305.14314
  37. ^Rafailov, R., et al. (2023). "Direct Preference Optimization: Your Language Model is Secretly a Reward Model." arxiv.org/...2305.18290
  38. ^Lee, H., et al. (2023). "RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback." arxiv.org/...2309.00267
  39. ^Su, J., et al. (2021). "RoFormer: Enhanced Transformer with Rotary Position Embedding." arxiv.org/...2104.09864
  40. ^Liu, Z., et al. (2024). "KAN: Kolmogorov-Arnold Networks." arxiv.org/...2404.19756
  41. ^Assran, M., et al. (2023). "Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture." openaccess.thecvf.com/...hitecture_CVPR_2023_paper
  42. ^Gu, A., and Dao, T. (2023). "Mamba: Linear-Time Sequence Modeling with Selective State Spaces." arxiv.org/...2312.00752
  43. ^Kim, M. J., et al. (2024). "OpenVLA: An Open-Source Vision-Language-Action Model." arxiv.org/...2406.09246
  44. ^DeepSeek-AI (2024). "DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model." arxiv.org/...2405.04434
  45. ^Ainslie, J., et al. (2023). "GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints." aclanthology.org/2023.emnlp-main.298
  46. ^Papineni, K., et al. (2002). "BLEU: a Method for Automatic Evaluation of Machine Translation." aclanthology.org/P02-1040
  47. ^Lin, C.-Y. (2004). "ROUGE: A Package for Automatic Evaluation of Summaries." aclanthology.org/W04-1013
  48. ^Banerjee, S., and Lavie, A. (2005). "METEOR: An Automatic Metric for MT Evaluation with Improved Correlation with Human Judgments." aclanthology.org/W05-0909
  49. ^Han, A. L. F., Wong, D. F., and Chao, L. S. (2012). "LEPOR: A Robust Evaluation Metric for Machine Translation with Augmented Factors." aclanthology.org/C12-2044
  50. ^Wang, A., et al. (2018). "GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding." openreview.net/forum
  51. ^Lopez-Paz, D., and Ranzato, M. (2017). "Gradient Episodic Memory for Continual Learning." papers.nips.cc/...ic-memory-for-continual-learning
  52. ^Kipf, T. N., and Welling, M. (2016). "Semi-Supervised Classification with Graph Convolutional Networks." arxiv.org/...1609.02907
  53. ^Yao, S., et al. (2023). "Tree of Thoughts: Deliberate Problem Solving with Large Language Models." arxiv.org/...2305.10601
  54. ^Model Context Protocol. "Specification." modelcontextprotocol.io/specification
  55. ^Eclipse Foundation. "SUMO: Simulation of Urban MObility." eclipse.dev/sumo
  56. ^Schubert, R., et al. (2024). "An open-source VISSIM calibration framework." pmc.ncbi.nlm.nih.gov/...PMC10878954
  57. ^Wielemaker, J., et al. (2010). "SWI-Prolog." arxiv.org/...1011.5332
  58. ^Google. "What is reCAPTCHA?" support.google.com/...6080904
  59. ^Bennett, C. H., and Brassard, G. (1984). "Quantum cryptography: Public key distribution and coin tossing." doi.org/...j.tcs.2014.05.025
  60. ^Alqahtani, A., et al. (2018). "ConditionaL Neural Networks." arxiv.org/...1804.02665
  61. ^Garnelo, M., et al. (2016). "Towards Deep Symbolic Reinforcement Learning." arxiv.org/...1609.05518
  62. ^Putzky, P., and Welling, M. (2017). "Recurrent Inference Machines for Solving Inverse Problems." arxiv.org/...1706.04008
  63. ^Kudo, T., and Richardson, J. (2018). "SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing." aclanthology.org/D18-2012
  64. ^McInnes, L., Healy, J., and Melville, J. (2018). "UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction." arxiv.org/...1802.03426
  65. ^Pennington, J., Socher, R., and Manning, C. D. (2014). "GloVe: Global Vectors for Word Representation." aclanthology.org/D14-1162
  66. ^Blei, D. M., Ng, A. Y., and Jordan, M. I. (2003). "Latent Dirichlet Allocation." jmlr.org/...blei03a
  67. ^Sennrich, R., Haddow, B., and Birch, A. (2016). "Neural Machine Translation of Rare Words with Subword Units." aclanthology.org/P16-1162
  68. ^Vaswani, A., et al. (2017). "Attention Is All You Need." arxiv.org/...1706.03762
  69. ^Ouyang, L., et al. (2022). "Training language models to follow instructions with human feedback." arxiv.org/...2203.02155
  70. ^Hugging Face. "Parameter-efficient fine-tuning." huggingface.co/...peft
  71. ^Hugging Face. "Supervised Fine-tuning Trainer." huggingface.co/...sft_trainer
  72. ^Wiggins, G. A., et al. (2024). "Automatic Music Transcription: A Survey." arxiv.org/...2406.15249
  73. ^Koller, O. (2023). "Quantitative survey of the state of the art in sign language recognition." link.springer.com/...s10209-023-00992-1
  74. ^Google Cloud. "What is BigQuery ML?" cloud.google.com/...bqml-introduction
  75. ^He, H., et al. (2008). "ADASYN: Adaptive Synthetic Sampling Approach for Imbalanced Learning." doi.org/...IJCNN.2008.4633969
  76. ^Brock, A., De, S., Smith, S. L., and Simonyan, K. (2021). "High-Performance Large-Scale Image Recognition Without Normalization." arxiv.org/...2102.06171
  77. ^Erhan, D., et al. (2010). "Why Does Unsupervised Pre-training Help Deep Learning?" proceedings.mlr.press/...erhan10a
  78. ^Salakhutdinov, R., and Mnih, A. (2008). "Bayesian Probabilistic Matrix Factorization Using Markov Chain Monte Carlo." cs.toronto.edu/...bpmf.pdf
  79. ^von Ahn, L., et al. (2003). "CAPTCHA: Using Hard AI Problems for Security." doi.org/...966389.966390
  80. ^Prokhorenkova, L., et al. (2018). "CatBoost: unbiased boosting with categorical features." proceedings.neurips.cc/...c41c24863285549-Abstract
  81. ^Bagnall, A., et al. (2017). "The great time series classification bake off." arxiv.org/...1602.01711
  82. ^Wang, C., and Blei, D. M. (2011). "Collaborative Topic Modeling for Recommending Scientific Articles." dl.acm.org/...2020408.2020480
  83. ^Cycorp. "The Cyc Knowledge Base." cyc.com/...cyc-knowledge-base
  84. ^Ester, M., et al. (1996). "A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with Noise." dl.acm.org/...3001460.3001507
  85. ^Petersen, B. K., et al. (2019). "Deep symbolic regression: Recovering mathematical expressions from data via risk-seeking policy gradients." arxiv.org/...1912.04871
  86. ^Montavon, G., et al. (2017). "Explaining nonlinear classification decisions with deep Taylor decomposition." doi.org/...j.patcog.2016.11.008
  87. ^Huang, G.-B., Zhu, Q.-Y., and Siew, C.-K. (2006). "Extreme learning machine: Theory and applications." doi.org/...j.neucom.2005.12.126
  88. ^Clevert, D.-A., Unterthiner, T., and Hochreiter, S. (2015). "Fast and Accurate Deep Network Learning by Exponential Linear Units." arxiv.org/...1511.07289
  89. ^Sun, Y., et al. (2019). "ERNIE: Enhanced Representation through Knowledge Integration." arxiv.org/...1904.09223
  90. ^Liu, Y., et al. (2018). "Secure Federated Transfer Learning." arxiv.org/...1812.03337
  91. ^van der Linden, I., Haned, H., and Kanoulas, E. (2019). "Global Aggregations of Local Explanations for Black Box models." arxiv.org/...1907.03039
  92. ^Mikolajczyk-Barela, A. (2023). "Data augmentation and explainability for bias discovery and mitigation in deep learning." arxiv.org/...2308.09464
  93. ^Hinton, G. (2021). "How to represent part-whole hierarchies in a neural network." arxiv.org/...2102.12627
  94. ^Cho, K., et al. (2014). "Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation." arxiv.org/...1406.1078
  95. ^Teh, Y. W., et al. (2004). "Sharing Clusters among Related Groups: Hierarchical Dirichlet Processes." proceedings.neurips.cc/...9739320d7a51c02-Abstract
  96. ^Hugging Face. "Documentation." huggingface.co/docs
  97. ^Sundararajan, M., Taly, A., and Yan, Q. (2017). "Axiomatic Attribution for Deep Networks." proceedings.mlr.press/...sundararajan17a
  98. ^Northcutt, C. G., Jiang, L., and Chuang, I. L. (2021). "Confident Learning: Estimating Uncertainty in Dataset Labels." jair.org/...12125
  99. ^Chen, T., et al. (2020). "A Simple Framework for Contrastive Learning of Visual Representations." proceedings.mlr.press/...chen20j
  100. ^Law, M., Russo, A., and Broda, K. (2021). "ILASP in the Fast Lane." ijcai.org/...0223.pdf
  101. ^Yeh, C.-K., et al. (2019). "On the (In)fidelity and Sensitivity for Explanations." arxiv.org/...1901.09392
  102. ^Chen, W., et al. (2024). "Internet of Agents: Weaving a Web of Heterogeneous Agents for Collaborative Intelligence." arxiv.org/...2407.07061
  103. ^Schuhmann, C., et al. (2022). "LAION-5B: An open large-scale dataset for training next generation image-text models." arxiv.org/...2210.08402
  104. ^Ke, G., et al. (2017). "LightGBM: A Highly Efficient Gradient Boosting Decision Tree." proceedings.neurips.cc/...669bdd9eb6b76fa-Abstract
  105. ^Ribeiro, M. T., Singh, S., and Guestrin, C. (2016). "Why Should I Trust You? Explaining the Predictions of Any Classifier." arxiv.org/...1602.04938
  106. ^Weinberger, K. Q., and Saul, L. K. (2009). "Distance Metric Learning for Large Margin Nearest Neighbor Classification." jmlr.org/...weinberger09a
  107. ^Germain, M., et al. (2015). "MADE: Masked Autoencoder for Distribution Estimation." proceedings.mlr.press/...germain15
  108. ^Papamakarios, G., Pavlakou, T., and Murray, I. (2017). "Masked Autoregressive Flow for Density Estimation." papers.nips.cc/...sive-flow-for-density-estimation
  109. ^Wang, J., et al. (2024). "Mixture-of-Agents Enhances Large Language Model Capabilities." arxiv.org/...2406.04692
  110. ^Stanley, K. O., and Miikkulainen, R. (2002). "Evolving Neural Networks through Augmenting Topologies." nn.cs.utexas.edu/...stanley.ec02.pdf
  111. ^NEST Initiative. "NEST Simulator documentation." nest-simulator.org
  112. ^Deb, K., et al. (2002). "A fast and elitist multiobjective genetic algorithm: NSGA-II." doi.org/...4235.996017
  113. ^Graves, A., Wayne, G., and Danihelka, I. (2014). "Neural Turing Machines." arxiv.org/...1410.5401
  114. ^Zhang, J., et al. (2020). "PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization." proceedings.mlr.press/...zhang20ae
  115. ^He, K., et al. (2015). "Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification." openaccess.thecvf.com/...Deep_into_ICCV_2015_paper
  116. ^Elkan, C., and Noto, K. (2008). "Learning classifiers from only positive and unlabeled data." cseweb.ucsd.edu/...posonly.pdf
  117. ^Petsiuk, V., Das, A., and Saenko, K. (2018). "RISE: Randomized Input Sampling for Explanation of Black-box Models." arxiv.org/...1806.07421
  118. ^Liu, Y., et al. (2019). "RoBERTa: A Robustly Optimized BERT Pretraining Approach." arxiv.org/...1907.11692
  119. ^Open Robotics. "ROS Documentation." docs.ros.org
  120. ^Yudkowsky, E. (2007). "Artificial Intelligence as a Positive and Negative Factor in Global Risk." intelligence.org/...AIPosNegFactor.pdf
  121. ^Chawla, N. V., et al. (2002). "SMOTE: Synthetic Minority Over-sampling Technique." jair.org/...10302
  122. ^Lapuschkin, S., et al. (2019). "Unmasking Clever Hans Predictors and Assessing What Machines Really Learn." arxiv.org/...1902.10178
  123. ^Zelikman, E., et al. (2022). "STaR: Bootstrapping Reasoning With Reasoning." arxiv.org/...2203.14465
  124. ^Raina, R., et al. (2007). "Self-taught learning: transfer learning from unlabeled data." dl.acm.org/...1273496.1273592
  125. ^Hensel, F., et al. (2021). "A Survey of Topological Machine Learning Methods." arxiv.org/...2101.05778
  126. ^van der Maaten, L., and Hinton, G. (2008). "Visualizing Data using t-SNE." jmlr.org/...vandermaaten08a
  127. ^Xie, Q., et al. (2019). "Unsupervised Data Augmentation for Consistency Training." arxiv.org/...1904.12848
  128. ^Serban, I. V., et al. (2016). "A Hierarchical Latent Variable Encoder-Decoder Model for Generating Dialogues." arxiv.org/...1605.06069
  129. ^Chen, T., and Guestrin, C. (2016). "XGBoost: A Scalable Tree Boosting System." doi.org/...2939672.2939785
  130. ^IBM. "CPLEX Optimizer documentation." ibm.com/...icos
  131. ^Trivedi, H., et al. (2022). "Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions." arxiv.org/...2212.10509
  132. ^Guu, K., et al. (2020). "REALM: Retrieval-Augmented Language Model Pre-Training." arxiv.org/...2002.08909
  133. ^Borgeaud, S., et al. (2021). "Improving language models by retrieving from trillions of tokens." arxiv.org/...2112.04426
  134. ^Tan, M., Pang, R., and Le, Q. V. (2020). "EfficientDet: Scalable and Efficient Object Detection." openaccess.thecvf.com/...Detection_CVPR_2020_paper
  135. ^Scikit-learn developers. "Linear and Quadratic Discriminant Analysis." scikit-learn.org/...lda_qda
  136. ^van den Oord, A., Vinyals, O., and Kavukcuoglu, K. (2017). "Neural Discrete Representation Learning." arxiv.org/...1711.00937
  137. ^Mirza, M., and Osindero, S. (2014). "Conditional Generative Adversarial Nets." arxiv.org/...1411.1784
  138. ^Radford, A., Metz, L., and Chintala, S. (2015). "Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks." arxiv.org/...1511.06434
  139. ^Middlehurst, M., Large, J., and Bagnall, A. (2021). "The Temporal Dictionary Ensemble (TDE) Classifier for Time Series Classification." arxiv.org/...2105.03841
  140. ^Vincent, P., et al. (2010). "Stacked Denoising Autoencoders: Learning Useful Representations in a Deep Network with a Local Denoising Criterion." jmlr.csail.mit.edu/...vincent10a
  141. ^Milletari, F., Navab, N., and Ahmadi, S.-A. (2016). "V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation." arxiv.org/...1606.04797

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

3 revisions · v4 · 6,644 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Cite this page: AI Wiki. "Acronyms." aiwiki.ai, updated 28 Jul 2026. CC BY 4.0. https://aiwiki.ai/wiki/acronyms

Suggest edit