🧠 NLP Concepts Glossary

← Back to Timeline
Era 1 · Symbolic Era 2 · Statistical Era 3 · ML & Structured Era 4 · Deep Learning Era 5 · Transformers Era 6 · Agentic AI
1

Era 1: 1950s–1970s: Symbolic & Rule-Based Era

AI Winter

Period of reduced funding and interest in AI research following unmet expectations.

Era 1

ALPAC Report

1966 report concluding MT was slower and more expensive than human translation; triggered first AI winter.

Era 1

Augmented Transition Network (ATN)

Extended finite-state automaton with registers and recursion for natural language parsing.

Era 1

Blocks World

Simplified domain of colored blocks used in early AI reasoning demonstrations.

Era 1

CYK Parsing Algorithm

O(n³) dynamic programming algorithm for CFG recognition and parsing.

Era 1

Chomsky Hierarchy

Formal classification of grammars into four types (Type-0 to Type-3) based on generative power.

Era 1

Conceptual Dependency

Schank's language-independent semantic representation using primitive actions.

Era 1

Context-Free Grammar (CFG)

Grammar formalism where production rules have a single non-terminal on the left-hand side.

Era 1

Definite Clause Grammar (DCG)

Logic-programming grammar formalism where rules are Horn clauses.

Era 1

ELIZA

Weizenbaum's pattern-matching chatbot simulating a Rogerian psychotherapist.

Era 1

Earley Parsing Algorithm

Efficient general CFG parsing; linear for practical grammars.

Era 1

Finite-State Automata

Abstract machine with a finite number of states; used for morphological analysis.

Era 1

First-Order Logic

Predicate logic with quantifiers (∀, ∃) for expressing complex relationships.

Era 1

Frame Theory

Minsky's data structure for stereotyped situations with slots, fillers, and defaults.

Era 1

Knowledge Representation

Formal structures for encoding facts, rules, and relationships about the world.

Era 1

LISP Programming Language

List-processing language designed for symbolic AI and recursive computation.

Era 1

LUNAR System

Woods' practical QA system accessing geological databases via ATN parsing.

Era 1

Logical Reasoning

Deductive inference using formal logic to derive conclusions from premises.

Era 1

Machine Translation (MT)

Automatic translation of text from one natural language to another.

Era 1

Micro-Planner

Planner used in SHRDLU for goal-directed reasoning about block manipulations.

Era 1

Montague Grammar

Model-theoretic semantics applying formal logic to natural language fragments.

Era 1

Pattern Matching

Template-based technique for recognizing and generating language patterns.

Era 1

Procedural Semantics

Semantic interpretation as computational procedure rather than static meaning.

Era 1

Prolog Programming Language

Logic-programming language based on Horn clauses and unification.

Era 1

Recursive Transition Network (RTN)

Network of states with recursive calls; equivalent in power to CFGs.

Era 1

SHRDLU

Winograd's blocks-world system integrating syntax, semantics, and reasoning.

Era 1

SIR System

Raphael's Semantic Information Retrieval using deductive inference networks.

Era 1

STUDENT System

Bobrow's system for solving algebra word problems via restricted English parsing.

Era 1

Semantic Network

Graph-based knowledge representation where concepts are nodes and relations are edges.

Era 1

Template-Based Generation

Surface realization via predefined sentence templates with slot filling.

Era 1

Theorem Proving

Automated derivation of mathematical proofs from axioms and inference rules.

Era 1

Transformational-Generative Grammar

Chomsky's theory that syntax is generated by phrase-structure rules plus transformations.

Era 1

Turing Test

Behavioral criterion for machine intelligence proposed by Alan Turing in 1950.

Era 1
2

Era 2: 1980s–1990s: Statistical Revolution

Bag-of-Words Representation

Text representation ignoring word order; frequency-based feature vector.

Era 2

British National Corpus (BNC)

100M-word corpus of modern British English (90% written, 10% spoken).

Era 2

Data Sparsity

Problem of zero counts for unseen events in statistical models.

Era 2

Data-Driven Probabilistic Modeling

Learning patterns from annotated corpora rather than handcrafting rules.

Era 2

Decision Tree

Tree-structured classifier learned from annotated training data.

Era 2

Document Retrieval

Finding relevant documents from a collection given a query.

Era 2

Expectation-Maximization (EM)

Iterative algorithm for maximum likelihood with hidden variables.

Era 2

Hidden Markov Model (HMM)

Probabilistic finite-state machine with hidden states and observable emissions.

Era 2

Information Extraction (IE)

Automatically extracting structured information from unstructured text.

Era 2

Language as Probability Distribution

Paradigm viewing language as a stochastic process rather than deterministic logic.

Era 2

Maximum Entropy Model (MaxEnt)

Discriminative framework maximizing entropy subject to feature constraints.

Era 2

Message Understanding Conference (MUC)

DARPA-sponsored IE evaluation series; established precision/recall/F1.

Era 2

N-gram Language Model

Markov model where word probability depends only on previous n−1 words.

Era 2

Naive Bayes Classifier

Probabilistic classifier assuming feature independence given the class.

Era 2

Noisy-Channel Model

Information-theoretic framework: source → noisy channel → observation.

Era 2

Part-of-Speech (POS) Tagging

Assigning grammatical categories (noun, verb, adjective) to each word.

Era 2

Penn Treebank

1M-word corpus of syntactically annotated Wall Street Journal text.

Era 2

Precision / Recall / F1

Standard information retrieval metrics: P = TP/(TP+FP), R = TP/(TP+FN), F1 = 2PR/(P+R).

Era 2

Shared Task Evaluation

Competitive evaluation with standardized data splits, metrics, and leaderboards.

Era 2

Smoothing (Good-Turing, Kneser-Ney)

Techniques to redistribute probability mass from seen to unseen events.

Era 2

Statistical Machine Translation (SMT)

Translation approach modeling P(target|source) via the noisy-channel framework.

Era 2

TF-IDF Weighting

Term frequency–inverse document frequency for scoring term importance.

Era 2

Text REtrieval Conference (TREC)

NIST-sponsored IR evaluation with large-scale test collections.

Era 2

Vector Space Model

Document representation as vectors in high-dimensional term space.

Era 2

WordNet

Lexical database of English with synonym sets and semantic relations.

Era 2
3

Era 3: 2000–2013: Machine Learning & Structured Prediction

Active Learning

Selective sampling of most informative instances for human annotation.

Era 3

BIO Tagging Scheme

Begin-Inside-Outside encoding for sequence labeling (NER, chunking).

Era 3

Chunking (Shallow Parsing)

Identifying non-recursive phrases (NP, VP, PP) without full parse trees.

Era 3

Co-Training

Training two classifiers on different feature views that teach each other.

Era 3

CoNLL Shared Tasks

Annual competitive evaluations standardizing data, metrics, and protocols.

Era 3

Conditional Random Field (CRF)

Discriminative undirected graphical model with global sequence normalization.

Era 3

Constituency Parsing (Berkeley Parser)

Phrase-structure parsing using PCFGs with latent annotations.

Era 3

Dependency Parsing

Syntactic analysis producing head-dependent relationships between words.

Era 3

Error Propagation in Pipelines

Compounding of errors across sequential processing stages.

Era 3

Feature Engineering

Manual design of lexical, syntactic, and orthographic feature templates.

Era 3

Generative vs. Discriminative Models

Trade-off between modeling P(X,Y) jointly vs. P(Y|X) conditionally.

Era 3

Graph-Based Parsing (MSTParser)

Parsing as finding maximum spanning tree over directed token graph.

Era 3

Kernel Methods

Technique mapping data to higher-dimensional spaces for linear separation.

Era 3

Label Bias Problem

Bias toward states with fewer outgoing transitions in locally normalized models.

Era 3

Labeled Attachment Score (LAS)

Dependency parsing metric counting correct head + label pairs.

Era 3

Latent Dirichlet Allocation (LDA)

Generative Bayesian model for discovering latent topics in document collections.

Era 3

MIRA (Margin Infused Relaxed Algorithm)

Online large-margin training for structured prediction.

Era 3

Named Entity Recognition (NER)

Identifying and classifying named entities (person, org, location) in text.

Era 3

Natural Language Toolkit (NLTK)

Open-source Python library for NLP education and research.

Era 3

PCFG with Latent Annotations (PCFG-LA)

Probabilistic CFG with automatically learned subcategories via EM.

Era 3

Self-Training

Using model predictions on unlabeled data as pseudo-labels for further training.

Era 3

Semi-Supervised Learning

Training with small labeled data plus large unlabeled data.

Era 3

Sentiment Analysis

Determining the emotional tone (positive, negative, neutral) of text.

Era 3

Structured Perceptron

Online learning algorithm updating weights based on entire output structure.

Era 3

Support Vector Machine (SVM)

Maximum-margin classifier with kernel methods for high-dimensional spaces.

Era 3

Topic Modeling

Unsupervised discovery of thematic structure in large text corpora.

Era 3

Transition-Based Parsing (MaltParser)

Deterministic, stack-based parsing with classifier-predicted actions.

Era 3

Unlabeled Attachment Score (UAS)

Dependency parsing metric counting correct head only.

Era 3

Word Clustering (Brown)

Hierarchical clustering of words based on bigram mutual information.

Era 3

scikit-learn

Python machine learning library with SVMs, decision trees, and evaluation tools.

Era 3
4

Era 4: 2014–2017: Deep Learning & Distributed Representations

Architecture Engineering

Designing neural architectures as the primary creative task.

Era 4

Bahdanau Attention

Additive attention scoring relevance of encoder states to decoder state.

Era 4

Bidirectional RNN

RNN processing sequence in both directions for richer context.

Era 4

Byte Pair Encoding (BPE)

Subword segmentation via iterative merging of frequent character pairs.

Era 4

Context-Independent Representations

Static embeddings assigning one vector per word type.

Era 4

Convolutional Neural Network (CNN) for Text

1D convolutions with multiple filter sizes for sentence classification.

Era 4

Dropout Regularization

Randomly masking units during training to prevent overfitting.

Era 4

Encoder-Decoder Architecture

Neural architecture with encoder (input representation) and decoder (output generation).

Era 4

End-to-End Training

Joint optimization of all model components via backpropagation.

Era 4

FastText (Subword Embeddings)

Word representations using character n-grams for OOV handling.

Era 4

Gated Recurrent Unit (GRU)

Simplified LSTM with update and reset gates; fewer parameters.

Era 4

GloVe (Global Vectors)

Log-bilinear model combining count-based and prediction-based approaches.

Era 4

Hierarchical Softmax

Tree-structured softmax reducing complexity from O(V) to O(log V).

Era 4

Long Short-Term Memory (LSTM)

RNN variant with forget/input/output gates solving vanishing gradients.

Era 4

Mini-Batch GPU Training

Parallel gradient computation on GPUs for scalable neural training.

Era 4

Multi-Head Attention

Multiple attention layers operating in parallel on different representation subspaces.

Era 4

Negative Sampling

Approximating softmax by sampling negative examples for training speed.

Era 4

Neural Attention Mechanism

Soft alignment allowing decoder to focus on relevant encoder states.

Era 4

Positional Encoding

Sinusoidal vectors injecting absolute/relative position information.

Era 4

Pretrained Word Vectors

Word embeddings trained on large corpora and reused across tasks.

Era 4

Residual Connections

Skip connections enabling gradient flow in very deep networks.

Era 4

Self-Attention

Direct pairwise interaction between all positions in a sequence.

Era 4

Semantic Similarity

Measuring relatedness via cosine similarity in vector space.

Era 4

Sequence-to-Sequence (Seq2Seq)

Encoder-decoder architecture mapping variable-length input to output.

Era 4

Span Extraction

Identifying answer start/end positions in a passage rather than classification.

Era 4

Static Embeddings

Word representations that do not change based on context.

Era 4

Subword Tokenization

Decomposing words into smaller units (characters, byte pairs) for OOV handling.

Era 4

Transfer Learning with Embeddings

Initializing models with pretrained embeddings for downstream tasks.

Era 4

Transformer Architecture

Attention-only architecture eliminating recurrence; O(1) sequential operations.

Era 4

Word Analogy Task

Evaluating embeddings by testing linear relationships (king − man + woman ≈ queen).

Era 4

Word Embeddings

Dense vector representations of words capturing semantic relationships.

Era 4

Word2Vec (Skip-gram / CBOW)

Efficient neural architectures for learning word embeddings from large corpora.

Era 4

WordPiece

Google's subword segmentation algorithm used in BERT.

Era 4
5

Era 5: 2018–2022: Pretrained Transformers & Emergent Capabilities

ALBERT (A Lite BERT)

Parameter-reduced BERT using factorized embeddings and cross-layer sharing.

Era 5

Adapters

Small bottleneck layers inserted between Transformer layers for task adaptation.

Era 5

Alignment

Ensuring AI systems act in accordance with human intent and values.

Era 5

Autoregressive Language Modeling

Predicting next token conditioned on all previous tokens.

Era 5

BIG-bench

Collaborative benchmark with 200+ diverse tasks probing LLM capabilities.

Era 5

Bias Amplification

Models encoding and amplifying societal biases from training data.

Era 5

Bidirectional Encoder Representations from Transformers (BERT)

Deep bidirectional Transformer pretrained with MLM and NSP.

Era 5

Catastrophic Forgetting

Degradation of previously learned skills during fine-tuning.

Era 5

Chain-of-Thought (CoT) Prompting

Prompting models to generate intermediate reasoning steps.

Era 5

Computational Inequality

Concentration of AI capability in well-funded organizations.

Era 5

DistilBERT

Smaller, faster BERT via knowledge distillation; 40% fewer parameters.

Era 5

ELECTRA

Pretraining by discriminating real vs. replaced tokens rather than masking.

Era 5

Emergent Capabilities

Abilities that appear suddenly and unpredictably at sufficient scale.

Era 5

Few-Shot Learning

Adapting to tasks with only a handful of examples in the prompt.

Era 5

Fine-Tuning

Adapting a pretrained model to a downstream task with labeled data.

Era 5

GLUE Benchmark

Multi-task NLU benchmark with 9 diverse tasks and aggregated score.

Era 5

Generative Pre-trained Transformer (GPT)

Decoder-only Transformer trained with left-to-right language modeling.

Era 5

Hallucination

Generation of plausible-sounding but factually incorrect content.

Era 5

Hugging Face Transformers

Open-source library providing pretrained models and training tools.

Era 5

In-Context Learning (ICL)

Task adaptation via prompts with examples; no parameter updates.

Era 5

Low-Rank Adaptation (LoRA)

Fine-tuning via low-rank decomposition of weight update matrices.

Era 5

MMLU (Massive Multitask Language Understanding)

Benchmark testing knowledge across 57 subjects.

Era 5

Masked Language Modeling (MLM)

Predicting randomly masked tokens from bidirectional context.

Era 5

Mixture of Experts (MoE)

Sparse activation routing tokens to subsets of experts.

Era 5

Next Sentence Prediction (NSP)

Binary classification of whether sentence B follows sentence A.

Era 5

Open Weights vs. API-Only

Trade-off between open model access and controlled API deployment.

Era 5

Parameter-Efficient Fine-Tuning

Updating small parameter subsets (adapters, LoRA) instead of full model.

Era 5

Prefix Tuning

Prepending trainable continuous vectors to keys and values in attention.

Era 5

Prompt Engineering

Designing input prompts to elicit desired behavior from LLMs.

Era 5

Proximal Policy Optimization (PPO)

RL algorithm optimizing policy while constraining deviation.

Era 5

Reinforcement Learning from Human Feedback (RLHF)

Aligning models with human preferences via reward modeling and PPO.

Era 5

Reward Model (RM)

Trained to score outputs by human preference rankings.

Era 5

RoBERTa (Robustly Optimized BERT)

BERT retrained with optimized hyperparameters and 10× more data.

Era 5

SQuAD (Stanford Question Answering Dataset)

Large-scale extractive reading comprehension benchmark.

Era 5

Scaling Laws

Power-law relationships between model size, data, compute, and performance.

Era 5

Span Corruption

Masking contiguous text spans and learning to reconstruct them (T5).

Era 5

SuperGLUE Benchmark

Harder successor to GLUE with more challenging tasks and minimal training data.

Era 5

Supervised Fine-Tuning (SFT)

Fine-tuning on human demonstrations as the first RLHF stage.

Era 5

Switch Transformers

Simplified MoE with single expert per token; scaled to 1.6T parameters.

Era 5

Text-to-Text Transfer Transformer (T5)

Encoder-decoder model casting all NLP tasks as text generation.

Era 5

Zero-Shot Learning

Performing tasks without any task-specific training examples.

Era 5
6

Era 6: 2023–Present: Agentic AI, Reasoning & Multimodal Systems

1M+ Token Context Window

Extended attention span enabling analysis of entire books or long videos.

Era 6

API-Only Models

Proprietary models accessible only via API (GPT-4, Claude).

Era 6

Agentic AI

AI systems capable of autonomous reasoning, planning, and action.

Era 6

Agentic Gap

Performance difference between AI agents and humans on real-world tasks.

Era 6

Alignment Complexity

Difficulty ensuring faithful instruction following with tool-augmented systems.

Era 6

Benchmark Gaming

Exploiting dataset artifacts rather than demonstrating genuine capability.

Era 6

CLIP (Contrastive Language-Image Pre-training)

Joint embedding space for text and images via contrastive learning.

Era 6

Chain-of-Thought (CoT)

Generating intermediate reasoning steps before final answers.

Era 6

Constitutional AI

Self-supervised alignment where AI critiques outputs against principles.

Era 6

Dynamic Tool Selection

Agent choosing which tools to invoke based on task requirements.

Era 6

Hallucination in RAG

Generation of false content even when retrieval sources are provided.

Era 6

Interactive Environment Evaluation

Testing agents in dynamic environments rather than static datasets.

Era 6

Large Language Model (LLM) as Agent

Using LLMs as the core reasoning engine for autonomous agents.

Era 6

Long-Context Modeling

Processing sequences of 1M+ tokens for long-document and video analysis.

Era 6

MMMU (Multimodal Understanding Benchmark)

College-level multimodal reasoning across 6 disciplines.

Era 6

Memory (Short/Long-Term)

Storing and retrieving information across agent interactions.

Era 6

Mind2Web (Generalist Web Agent)

Dataset for generalist web agents across 137 websites.

Era 6

Mixture of Experts (MoE)

Sparse activation routing tokens to subsets of experts.

Era 6

Multi-Agent Systems

Multiple specialized agents collaborating on complex tasks.

Era 6

Multimodal LLM

Model processing and generating text, images, audio, and video.

Era 6

Native Multimodality

Unified architecture processing all modalities without separate modules.

Era 6

Open-Source LLM Ecosystem

Community-driven development around open-weight models (LLaMA, Mixtral).

Era 6

Planning and Reasoning

Decomposing goals into sub-goals and executing them sequentially.

Era 6

Prompt Injection

Attack where malicious input via tool outputs hijacks model behavior.

Era 6

Quantized Low-Rank Adaptation (QLoRA)

LoRA with 4-bit quantization for memory-efficient fine-tuning.

Era 6

ReAct (Reasoning + Acting)

Interleaving reasoning traces with executable actions in dynamic environments.

Era 6

Real-Time Interaction

Low-latency audio-visual interaction with sub-second response times.

Era 6

Reflexion (Self-Reflective Agents)

Agents with verbal self-critique and dynamic memory for error correction.

Era 6

Reinforcement Learning from AI Feedback (RLAIF)

Using AI-generated preferences instead of human feedback for alignment.

Era 6

Retrieval-Augmented Generation (RAG)

Combining parametric memory with external retrieval for grounded generation.

Era 6

SWE-bench (Software Engineering Benchmark)

Real GitHub issue resolution benchmark for coding agents.

Era 6

Safety in Agentic Systems

Ensuring autonomous agents do not cause harm through their actions.

Era 6

Self-Improvement

Models improving through iterative self-critique without human annotation.

Era 6

Sparse Mixture of Experts (SMoE)

MoE architecture where only a fraction of parameters are active per token.

Era 6

Speculative Decoding

Draft model predicts tokens; target model verifies for inference speedup.

Era 6

Tool Failure Cascade

Compounding errors when API calls fail in multi-step agent trajectories.

Era 6

Tool Use / Function Calling

Models invoking external APIs to extend beyond parametric knowledge.

Era 6

ToolFormer (Self-Supervised Tool Learning)

Self-supervised method for teaching models to use tools.

Era 6

Tree-of-Thoughts

Exploring multiple reasoning paths and backtracking from dead ends.

Era 6

Vision Transformer (ViT)

Transformer architecture applied to image patches for visual understanding.

Era 6

WebArena (Web Navigation Benchmark)

Self-contained web environment for evaluating autonomous web agents.

Era 6

Weight-Decomposed Low-Rank Adaptation (DoRA)

Decomposing weights into magnitude and direction for fine-tuning.

Era 6