Period of reduced funding and interest in AI research following unmet expectations.
Era 11966 report concluding MT was slower and more expensive than human translation; triggered first AI winter.
Era 1Extended finite-state automaton with registers and recursion for natural language parsing.
Era 1Simplified domain of colored blocks used in early AI reasoning demonstrations.
Era 1O(n³) dynamic programming algorithm for CFG recognition and parsing.
Era 1Formal classification of grammars into four types (Type-0 to Type-3) based on generative power.
Era 1Schank's language-independent semantic representation using primitive actions.
Era 1Grammar formalism where production rules have a single non-terminal on the left-hand side.
Era 1Logic-programming grammar formalism where rules are Horn clauses.
Era 1Weizenbaum's pattern-matching chatbot simulating a Rogerian psychotherapist.
Era 1Efficient general CFG parsing; linear for practical grammars.
Era 1Abstract machine with a finite number of states; used for morphological analysis.
Era 1Predicate logic with quantifiers (∀, ∃) for expressing complex relationships.
Era 1Minsky's data structure for stereotyped situations with slots, fillers, and defaults.
Era 1Formal structures for encoding facts, rules, and relationships about the world.
Era 1List-processing language designed for symbolic AI and recursive computation.
Era 1Woods' practical QA system accessing geological databases via ATN parsing.
Era 1Deductive inference using formal logic to derive conclusions from premises.
Era 1Automatic translation of text from one natural language to another.
Era 1Planner used in SHRDLU for goal-directed reasoning about block manipulations.
Era 1Model-theoretic semantics applying formal logic to natural language fragments.
Era 1Template-based technique for recognizing and generating language patterns.
Era 1Semantic interpretation as computational procedure rather than static meaning.
Era 1Logic-programming language based on Horn clauses and unification.
Era 1Network of states with recursive calls; equivalent in power to CFGs.
Era 1Winograd's blocks-world system integrating syntax, semantics, and reasoning.
Era 1Raphael's Semantic Information Retrieval using deductive inference networks.
Era 1Bobrow's system for solving algebra word problems via restricted English parsing.
Era 1Graph-based knowledge representation where concepts are nodes and relations are edges.
Era 1Surface realization via predefined sentence templates with slot filling.
Era 1Automated derivation of mathematical proofs from axioms and inference rules.
Era 1Chomsky's theory that syntax is generated by phrase-structure rules plus transformations.
Era 1Behavioral criterion for machine intelligence proposed by Alan Turing in 1950.
Era 1Text representation ignoring word order; frequency-based feature vector.
Era 2100M-word corpus of modern British English (90% written, 10% spoken).
Era 2Problem of zero counts for unseen events in statistical models.
Era 2Learning patterns from annotated corpora rather than handcrafting rules.
Era 2Tree-structured classifier learned from annotated training data.
Era 2Finding relevant documents from a collection given a query.
Era 2Iterative algorithm for maximum likelihood with hidden variables.
Era 2Probabilistic finite-state machine with hidden states and observable emissions.
Era 2Automatically extracting structured information from unstructured text.
Era 2Paradigm viewing language as a stochastic process rather than deterministic logic.
Era 2Discriminative framework maximizing entropy subject to feature constraints.
Era 2DARPA-sponsored IE evaluation series; established precision/recall/F1.
Era 2Markov model where word probability depends only on previous n−1 words.
Era 2Probabilistic classifier assuming feature independence given the class.
Era 2Information-theoretic framework: source → noisy channel → observation.
Era 2Assigning grammatical categories (noun, verb, adjective) to each word.
Era 21M-word corpus of syntactically annotated Wall Street Journal text.
Era 2Standard information retrieval metrics: P = TP/(TP+FP), R = TP/(TP+FN), F1 = 2PR/(P+R).
Era 2Competitive evaluation with standardized data splits, metrics, and leaderboards.
Era 2Techniques to redistribute probability mass from seen to unseen events.
Era 2Translation approach modeling P(target|source) via the noisy-channel framework.
Era 2Term frequency–inverse document frequency for scoring term importance.
Era 2NIST-sponsored IR evaluation with large-scale test collections.
Era 2Document representation as vectors in high-dimensional term space.
Era 2Lexical database of English with synonym sets and semantic relations.
Era 2Selective sampling of most informative instances for human annotation.
Era 3Begin-Inside-Outside encoding for sequence labeling (NER, chunking).
Era 3Identifying non-recursive phrases (NP, VP, PP) without full parse trees.
Era 3Training two classifiers on different feature views that teach each other.
Era 3Annual competitive evaluations standardizing data, metrics, and protocols.
Era 3Discriminative undirected graphical model with global sequence normalization.
Era 3Phrase-structure parsing using PCFGs with latent annotations.
Era 3Syntactic analysis producing head-dependent relationships between words.
Era 3Compounding of errors across sequential processing stages.
Era 3Manual design of lexical, syntactic, and orthographic feature templates.
Era 3Trade-off between modeling P(X,Y) jointly vs. P(Y|X) conditionally.
Era 3Parsing as finding maximum spanning tree over directed token graph.
Era 3Technique mapping data to higher-dimensional spaces for linear separation.
Era 3Bias toward states with fewer outgoing transitions in locally normalized models.
Era 3Dependency parsing metric counting correct head + label pairs.
Era 3Generative Bayesian model for discovering latent topics in document collections.
Era 3Online large-margin training for structured prediction.
Era 3Identifying and classifying named entities (person, org, location) in text.
Era 3Open-source Python library for NLP education and research.
Era 3Probabilistic CFG with automatically learned subcategories via EM.
Era 3Using model predictions on unlabeled data as pseudo-labels for further training.
Era 3Training with small labeled data plus large unlabeled data.
Era 3Determining the emotional tone (positive, negative, neutral) of text.
Era 3Online learning algorithm updating weights based on entire output structure.
Era 3Maximum-margin classifier with kernel methods for high-dimensional spaces.
Era 3Unsupervised discovery of thematic structure in large text corpora.
Era 3Deterministic, stack-based parsing with classifier-predicted actions.
Era 3Dependency parsing metric counting correct head only.
Era 3Hierarchical clustering of words based on bigram mutual information.
Era 3Python machine learning library with SVMs, decision trees, and evaluation tools.
Era 3Designing neural architectures as the primary creative task.
Era 4Additive attention scoring relevance of encoder states to decoder state.
Era 4RNN processing sequence in both directions for richer context.
Era 4Subword segmentation via iterative merging of frequent character pairs.
Era 4Static embeddings assigning one vector per word type.
Era 41D convolutions with multiple filter sizes for sentence classification.
Era 4Randomly masking units during training to prevent overfitting.
Era 4Neural architecture with encoder (input representation) and decoder (output generation).
Era 4Joint optimization of all model components via backpropagation.
Era 4Word representations using character n-grams for OOV handling.
Era 4Simplified LSTM with update and reset gates; fewer parameters.
Era 4Log-bilinear model combining count-based and prediction-based approaches.
Era 4Tree-structured softmax reducing complexity from O(V) to O(log V).
Era 4RNN variant with forget/input/output gates solving vanishing gradients.
Era 4Parallel gradient computation on GPUs for scalable neural training.
Era 4Multiple attention layers operating in parallel on different representation subspaces.
Era 4Approximating softmax by sampling negative examples for training speed.
Era 4Soft alignment allowing decoder to focus on relevant encoder states.
Era 4Sinusoidal vectors injecting absolute/relative position information.
Era 4Word embeddings trained on large corpora and reused across tasks.
Era 4Skip connections enabling gradient flow in very deep networks.
Era 4Direct pairwise interaction between all positions in a sequence.
Era 4Measuring relatedness via cosine similarity in vector space.
Era 4Encoder-decoder architecture mapping variable-length input to output.
Era 4Identifying answer start/end positions in a passage rather than classification.
Era 4Word representations that do not change based on context.
Era 4Decomposing words into smaller units (characters, byte pairs) for OOV handling.
Era 4Initializing models with pretrained embeddings for downstream tasks.
Era 4Attention-only architecture eliminating recurrence; O(1) sequential operations.
Era 4Evaluating embeddings by testing linear relationships (king − man + woman ≈ queen).
Era 4Dense vector representations of words capturing semantic relationships.
Era 4Efficient neural architectures for learning word embeddings from large corpora.
Era 4Google's subword segmentation algorithm used in BERT.
Era 4Parameter-reduced BERT using factorized embeddings and cross-layer sharing.
Era 5Small bottleneck layers inserted between Transformer layers for task adaptation.
Era 5Ensuring AI systems act in accordance with human intent and values.
Era 5Predicting next token conditioned on all previous tokens.
Era 5Collaborative benchmark with 200+ diverse tasks probing LLM capabilities.
Era 5Models encoding and amplifying societal biases from training data.
Era 5Deep bidirectional Transformer pretrained with MLM and NSP.
Era 5Degradation of previously learned skills during fine-tuning.
Era 5Prompting models to generate intermediate reasoning steps.
Era 5Concentration of AI capability in well-funded organizations.
Era 5Smaller, faster BERT via knowledge distillation; 40% fewer parameters.
Era 5Pretraining by discriminating real vs. replaced tokens rather than masking.
Era 5Abilities that appear suddenly and unpredictably at sufficient scale.
Era 5Adapting to tasks with only a handful of examples in the prompt.
Era 5Adapting a pretrained model to a downstream task with labeled data.
Era 5Multi-task NLU benchmark with 9 diverse tasks and aggregated score.
Era 5Decoder-only Transformer trained with left-to-right language modeling.
Era 5Generation of plausible-sounding but factually incorrect content.
Era 5Open-source library providing pretrained models and training tools.
Era 5Task adaptation via prompts with examples; no parameter updates.
Era 5Fine-tuning via low-rank decomposition of weight update matrices.
Era 5Benchmark testing knowledge across 57 subjects.
Era 5Predicting randomly masked tokens from bidirectional context.
Era 5Sparse activation routing tokens to subsets of experts.
Era 5Binary classification of whether sentence B follows sentence A.
Era 5Trade-off between open model access and controlled API deployment.
Era 5Updating small parameter subsets (adapters, LoRA) instead of full model.
Era 5Prepending trainable continuous vectors to keys and values in attention.
Era 5Designing input prompts to elicit desired behavior from LLMs.
Era 5RL algorithm optimizing policy while constraining deviation.
Era 5Aligning models with human preferences via reward modeling and PPO.
Era 5Trained to score outputs by human preference rankings.
Era 5BERT retrained with optimized hyperparameters and 10× more data.
Era 5Large-scale extractive reading comprehension benchmark.
Era 5Power-law relationships between model size, data, compute, and performance.
Era 5Masking contiguous text spans and learning to reconstruct them (T5).
Era 5Harder successor to GLUE with more challenging tasks and minimal training data.
Era 5Fine-tuning on human demonstrations as the first RLHF stage.
Era 5Simplified MoE with single expert per token; scaled to 1.6T parameters.
Era 5Encoder-decoder model casting all NLP tasks as text generation.
Era 5Performing tasks without any task-specific training examples.
Era 5Extended attention span enabling analysis of entire books or long videos.
Era 6Proprietary models accessible only via API (GPT-4, Claude).
Era 6AI systems capable of autonomous reasoning, planning, and action.
Era 6Performance difference between AI agents and humans on real-world tasks.
Era 6Difficulty ensuring faithful instruction following with tool-augmented systems.
Era 6Exploiting dataset artifacts rather than demonstrating genuine capability.
Era 6Joint embedding space for text and images via contrastive learning.
Era 6Generating intermediate reasoning steps before final answers.
Era 6Self-supervised alignment where AI critiques outputs against principles.
Era 6Agent choosing which tools to invoke based on task requirements.
Era 6Generation of false content even when retrieval sources are provided.
Era 6Testing agents in dynamic environments rather than static datasets.
Era 6Using LLMs as the core reasoning engine for autonomous agents.
Era 6Processing sequences of 1M+ tokens for long-document and video analysis.
Era 6College-level multimodal reasoning across 6 disciplines.
Era 6Storing and retrieving information across agent interactions.
Era 6Dataset for generalist web agents across 137 websites.
Era 6Sparse activation routing tokens to subsets of experts.
Era 6Multiple specialized agents collaborating on complex tasks.
Era 6Model processing and generating text, images, audio, and video.
Era 6Unified architecture processing all modalities without separate modules.
Era 6Community-driven development around open-weight models (LLaMA, Mixtral).
Era 6Decomposing goals into sub-goals and executing them sequentially.
Era 6Attack where malicious input via tool outputs hijacks model behavior.
Era 6LoRA with 4-bit quantization for memory-efficient fine-tuning.
Era 6Interleaving reasoning traces with executable actions in dynamic environments.
Era 6Low-latency audio-visual interaction with sub-second response times.
Era 6Agents with verbal self-critique and dynamic memory for error correction.
Era 6Using AI-generated preferences instead of human feedback for alignment.
Era 6Combining parametric memory with external retrieval for grounded generation.
Era 6Real GitHub issue resolution benchmark for coding agents.
Era 6Ensuring autonomous agents do not cause harm through their actions.
Era 6Models improving through iterative self-critique without human annotation.
Era 6MoE architecture where only a fraction of parameters are active per token.
Era 6Draft model predicts tokens; target model verifies for inference speedup.
Era 6Compounding errors when API calls fail in multi-step agent trajectories.
Era 6Models invoking external APIs to extend beyond parametric knowledge.
Era 6Self-supervised method for teaching models to use tools.
Era 6Exploring multiple reasoning paths and backtracking from dead ends.
Era 6Transformer architecture applied to image patches for visual understanding.
Era 6Self-contained web environment for evaluating autonomous web agents.
Era 6Decomposing weights into magnitude and direction for fine-tuning.
Era 6