Paper Breakdowns
Plain-language walkthroughs of the papers that shaped modern AI — the core idea, the results that mattered, and how it connects to what you already know.
3D Gaussian Splatting
3D Gaussian Splatting for Real-Time Radiance Field Rendering
The 2023 paper that dethroned NeRFs for novel view synthesis by replacing the neural network with millions of explicit 3D splats, achieving real-time 1080p rendering.
Read breakdownAdversarial Examples (FGSM)
Explaining and Harnessing Adversarial Examples
The 2014 paper by Goodfellow et al. that exposed a terrifying flaw in neural networks: adding invisible, mathematically calculated noise to an image causes the model to confidently misclassify it.
Read breakdownAlphaFold 2
Highly Accurate Protein Structure Prediction with AlphaFold
The 2021 DeepMind paper that solved a 50-year-old grand challenge in biology by using AI to predict the 3D structure of a protein from its 1D amino acid sequence.
Read breakdownAre Emergent Abilities a Mirage?
Are Emergent Abilities of Large Language Models a Mirage?
A powerful rebuttal arguing that the 'sudden' appearance of capabilities in LLMs is merely a statistical illusion caused by researchers choosing non-linear, discontinuous metrics.
Read breakdownAttention Is All You Need
Attention Is All You Need
The landmark 2017 paper that introduced the Transformer architecture, replacing RNNs with self-attention and launching the modern era of large language models.
Read breakdownAWQ
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Introduced Activation-aware Weight Quantization, a method that preserves LLM performance by identifying and protecting a tiny fraction of highly salient weights based on activation magnitudes.
Read breakdownBERT
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Introduced bidirectional pretraining by masked language modelling, setting a new standard for natural language understanding tasks.
Read breakdownBetter & Faster LLMs via Multi-token Prediction
Better & Faster Large Language Models via Multi-token Prediction
Proposed training LLMs to predict the next $N$ tokens simultaneously rather than just the single next token, improving reasoning capabilities and creating a built-in draft model for fast speculative decoding.
Read breakdownBillion-Scale Similarity Search with GPUs
Billion-Scale Similarity Search with GPUs
The 2017 paper by Facebook AI Research that introduced FAISS, making massive-scale vector similarity search practical.
Read breakdownBitNet (1-bit LLMs)
BitNet: Scaling 1-bit Transformers for Large Language Models
The 2023 Microsoft paper that challenged the fundamental math of neural networks, proving you can train LLMs where weights are just +1 or -1, eliminating matrix multiplication entirely.
Read breakdownChain-of-Thought Prompting
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
The 2022 Google Brain paper that unlocked complex reasoning in LLMs by simply prompting them to show their work step-by-step before answering.
Read breakdownChinchilla Scaling Laws
Training Compute-Optimal Large Language Models
The 2022 DeepMind paper that redefined how LLMs are scaled, proving that most existing models were too large and under-trained on too little data.
Read breakdownClassifier-Free Guidance (CFG)
Classifier-Free Diffusion Guidance
The 2022 paper that unlocked high-quality text-to-image generation by teaching diffusion models to heavily prioritize the text prompt over the unconditional image prior.
Read breakdownClaude 3
The Claude 3 Model Family: Opus, Sonnet, Haiku
The 2024 Anthropic paper detailing a family of models that pushed the frontier of AI capabilities, heavily utilizing Constitutional AI and Constitutional alignment.
Read breakdownCLIP
Learning Transferable Visual Models From Natural Language Supervision
The 2021 OpenAI paper that aligned text and images in a shared embedding space, unlocking zero-shot classification and the modern generative image era.
Read breakdownColBERT
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
The 2020 paper that introduced Late Interaction, bridging the speed of dual-encoders with the accuracy of cross-encoders for neural retrieval.
Read breakdownConcrete Problems in AI Safety
Concrete Problems in AI Safety
The 2016 paper that defined the modern field of AI Alignment by clearly categorizing how optimization processes can go disastrously wrong in the real world.
Read breakdownConformal Prediction
A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification
The 2021 tutorial that popularized Conformal Prediction, a mathematical framework to add rigorous uncertainty bounds (confidence intervals) to any machine learning model.
Read breakdownConsistency Models
Consistency Models
The 2023 paper from OpenAI that enables high-quality diffusion generation in just a single step, bypassing the slow iterative sampling process entirely.
Read breakdownConstitutional AI
Constitutional AI: Harmlessness from AI Feedback
The 2022 Anthropic paper that introduced a method to align models using AI-generated feedback rather than human labels, scaling alignment securely and transparently.
Read breakdownControlNet
Adding Conditional Control to Text-to-Image Diffusion Models
The 2023 paper that introduced a way to add precise spatial control (like edge maps or human poses) to large text-to-image diffusion models without retraining them.
Read breakdownConvNeXt
A ConvNet for the 2020s
The 2022 paper that modernized the standard ResNet architecture to prove CNNs could still match or beat Vision Transformers on accuracy and scalability.
Read breakdownData Extraction Attacks
Extracting Training Data from Large Language Models
The 2020 paper that exposed a critical privacy flaw in LLMs, proving that massive generative models verbatim memorize their training data, which can be easily extracted by users.
Read breakdownDDIM
Denoising Diffusion Implicit Models
The 2020 paper that dramatically accelerated diffusion model sampling, reducing generation time from thousands of steps to a few dozen without retraining.
Read breakdownDDPM
Denoising Diffusion Probabilistic Models
The 2020 paper that proved diffusion models could generate high-quality images by learning to reverse a gradual noising process.
Read breakdownDeep Double Descent
Deep Double Descent: Where Bigger Models and More Data Hurt
The 2019 paper that broke classical statistical theory, proving that making a neural network "too big" actually causes its error rate to drop again after an initial spike.
Read breakdownDeepSeek-R1
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
The 2025 breakthrough that proved 'aha' moments and advanced reasoning capabilities could emerge purely from large-scale reinforcement learning, without needing millions of human-written reasoning examples.
Read breakdownDeepSeek-V2
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
The 2024 paper that disrupted the AI pricing landscape by introducing Multi-Head Latent Attention (MLA), drastically reducing the memory required for the KV Cache.
Read breakdownDeepSeekMath (GRPO)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
The 2024 paper that introduced Group Relative Policy Optimization (GRPO), a highly efficient alternative to PPO that eliminates the need for a separate value model during reinforcement learning.
Read breakdownDense Passage Retrieval (DPR)
Dense Passage Retrieval for Open-Domain Question Answering
The 2020 paper that proved dense neural embeddings could outperform traditional lexical search (like BM25) for open-domain question answering.
Read breakdownDiffusion Transformers (DiT)
Scalable Diffusion Models with Transformers
The 2022 paper that proved Transformers could replace the U-Net as the backbone for diffusion models, unlocking predictable scaling laws for image generation.
Read breakdownDINOv2
DINOv2: Learning Robust Visual Features without Supervision
The 2023 paper from Meta that produced state-of-the-art self-supervised visual features, matching supervised models without using any labels or text.
Read breakdownDirect Preference Optimization (DPO)
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
The 2023 Stanford paper that revolutionized LLM alignment by mathematically proving you can skip the complex Reward Model and PPO phases of RLHF, aligning models directly from preference data.
Read breakdownDouble/Debiased ML
Double/Debiased Machine Learning for Treatment and Structural Parameters
The 2016 economics and statistics paper that proved how to use highly flexible machine learning models to infer true causal effects, without bias.
Read breakdownDP-SGD
Deep Learning with Differential Privacy
The 2016 Google paper that proved you can train deep neural networks while mathematically guaranteeing the privacy of the individuals in the training dataset.
Read breakdownEfficient Memory Management for LLM Serving
Efficient Memory Management for Large Language Model Serving with PagedAttention
Introduced PagedAttention and the vLLM engine, applying operating system paging concepts to the KV cache to eliminate memory fragmentation and drastically increase serving throughput.
Read breakdownEfficient Streaming LMs with Attention Sinks
Efficient Streaming Language Models with Attention Sinks
Discovered that LLMs naturally use the first few tokens as 'attention sinks' to dump excess attention scores, enabling a simple fix to keep models generating text infinitely without crashing.
Read breakdownEfficientNet
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
The 2019 paper that introduced compound scaling, a principled way to scale up convolutional networks across depth, width, and resolution.
Read breakdownEmergent Abilities of LLMs
Emergent Abilities of Large Language Models
Argued that certain complex capabilities in LLMs appear suddenly and unpredictably only after the model crosses a specific scale threshold, becoming a central tenet of the AI scaling hypothesis.
Read breakdownFast Inference via Speculative Decoding
Fast Inference from Transformers via Speculative Decoding
Introduced Speculative Decoding, a technique that drastically speeds up LLM inference by using a small, fast draft model to generate tokens, which a larger model then verifies in parallel.
Read breakdownFID (Fréchet Inception Distance)
GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
The 2017 paper that introduced Fréchet Inception Distance (FID), solving the critical problem of how to mathematically measure the quality of AI-generated images.
Read breakdownFlashAttention
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
The 2022 Stanford paper that rewrote the attention algorithm to be hardware-aware, drastically speeding up Transformers and unlocking massive context windows.
Read breakdownFlashAttention-2
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
An optimized iteration of FlashAttention that better partitions work across GPU thread blocks, reaching closer to the theoretical maximum speed of the hardware.
Read breakdownFlow Matching
Flow Matching for Generative Modeling
The 2022 paper that generalized diffusion models into a simpler, simulation-free framework based on Continuous Normalizing Flows and vector fields.
Read breakdownGemini
Gemini: A Family of Highly Capable Multimodal Models
The 2023 Google DeepMind paper introducing a natively multimodal model family built from the ground up to reason seamlessly across text, images, audio, and video.
Read breakdownGPT-4
GPT-4 Technical Report
The 2023 technical report from OpenAI detailing the model that defined the frontier of AI capabilities, demonstrating human-level performance on professional benchmarks.
Read breakdownGPTQ
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Introduced a highly efficient post-training quantization method that can compress 175B parameter models down to 3 or 4 bits per weight with negligible accuracy degradation.
Read breakdownGQA
GQA: Training Generalized Multi-Query Attention
Introduced Grouped-Query Attention, striking an optimal balance between the high quality of Multi-Head Attention and the fast inference speed of Multi-Query Attention.
Read breakdownHNSW
Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs
The 2016 paper that introduced the Navigable Small-World graph, the reigning standard for extremely fast approximate nearest neighbor search.
Read breakdownInstructGPT
Training language models to follow instructions with human feedback
The 2022 paper that detailed the RLHF methodology used to align GPT-3 into a helpful assistant, directly laying the groundwork for ChatGPT.
Read breakdownLanguage Models are Few-Shot Learners
Language Models are Few-Shot Learners
The GPT-3 paper that formalized in-context learning, showing that massive scale allows models to learn new tasks simply from examples in the prompt.
Read breakdownLanguage Models are Unsupervised Multitask Learners
Language Models are Unsupervised Multitask Learners
The paper introducing GPT-2, demonstrating that scaling up a language model allows it to perform various downstream tasks zero-shot without explicit fine-tuning.
Read breakdownLatent Diffusion Models (LDM)
High-Resolution Image Synthesis with Latent Diffusion Models
The 2021 paper that brought diffusion models to the masses by running the generative process in a compressed latent space, drastically reducing compute requirements.
Read breakdownLet's Verify Step by Step
Let's Verify Step by Step
A 2023 OpenAI paper that showed training reward models to evaluate every single step of a reasoning chain (Process Supervision) drastically outperforms models trained only to evaluate the final answer (Outcome Supervision).
Read breakdownLIME
Why Should I Trust You?: Explaining the Predictions of Any Classifier
The 2016 paper that introduced Local Interpretable Model-agnostic Explanations, a method to figure out exactly why a "black box" AI made a specific decision.
Read breakdownLLaMA 1
LLaMA: Open and Efficient Foundation Language Models
The 2023 paper from Meta that introduced the first highly capable, open-weights foundation model, sparking the open-source LLM revolution.
Read breakdownLlama 2
Llama 2: Open Foundation and Fine-Tuned Chat Models
The 2023 paper from Meta that introduced a commercially viable, chat-tuned foundation model, detailing the immense effort required for RLHF and safety alignment.
Read breakdownLlama 3
The Llama 3 Herd of Models
The 2024 Meta technical report detailing a massive scale-up in data and compute, proving that dense models can reach frontier-level capabilities through sheer scale and data quality.
Read breakdownLLaVA
Visual Instruction Tuning
The 2023 paper that demonstrated how to turn an open-source text LLM into a powerful multimodal model by projecting visual features into the text token space.
Read breakdownLoRA
LoRA: Low-Rank Adaptation of Large Language Models
The 2021 Microsoft paper that made fine-tuning massive models accessible to anyone by training only a tiny fraction of the parameters.
Read breakdownLoRA
LoRA: Low-Rank Adaptation of Large Language Models
The 2021 paper from Microsoft that introduced Low-Rank Adaptation, allowing massive models to be fine-tuned quickly and cheaply by only training a tiny fraction of the parameters.
Read breakdownLost in the Middle
Lost in the Middle: How Language Models Use Long Contexts
A landmark empirical study revealing that despite massive context windows, LLMs systematically fail to retrieve information hidden in the middle of long documents.
Read breakdownLottery Ticket Hypothesis
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
The 2018 MIT paper that proved massive neural networks contain tiny, sparse subnetworks that can learn just as fast and achieve the exact same accuracy as the giant model.
Read breakdownMamba
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Introduced a new Selective State Space Model architecture that achieves Transformer-level quality with linear-time inference and hardware-aware scaling.
Read breakdownMatryoshka Representation Learning
Matryoshka Representation Learning
The 2022 paper that introduced a way to train embedding models so their output vectors can be truncated to smaller sizes without retraining.
Read breakdownMegatron-LM
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Introduced an elegant technique for Tensor Parallelism, allowing massive Transformer models to be split across multiple GPUs by partitioning the matrix multiplications themselves.
Read breakdownMixtral 8x7B
Mixtral of Experts
The 2024 paper from Mistral AI that brought Sparse Mixture-of-Experts (MoE) to the open-source community, enabling a 47B parameter model to run at the speed of a 14B model.
Read breakdownModel Cards
Model Cards for Model Reporting
The 2018 paper that established the industry standard for AI transparency, proposing that every machine learning model should be accompanied by a 'nutrition label' detailing its performance, limits, and biases.
Read breakdownNeRF
NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis
The 2020 paper that revolutionized 3D computer vision by proving a neural network could "memorize" a 3D scene and render photorealistic novel views.
Read breakdownOrca
Orca: Progressive Learning from Complex Explanation Traces of GPT-4
The 2023 Microsoft paper that improved upon synthetic data distillation by teaching the smaller model the *reasoning process* of the teacher model, not just the final answer.
Read breakdownPAC Learning
A Theory of the Learnable
The 1984 paper that mathematically defined what it actually means for a computer to 'learn', laying the foundational theory for modern machine learning.
Read breakdownPhi-1
Textbooks Are All You Need
The 2023 Microsoft paper that proved massive parameter counts aren't necessary if the training data is of extremely high, 'textbook' quality.
Read breakdownQLoRA
QLoRA: Efficient Finetuning of Quantized LLMs
The 2023 paper that combined 4-bit quantization with LoRA, proving you could fine-tune a massive 65B parameter model on a single consumer GPU without losing performance.
Read breakdownQLoRA
QLoRA: Efficient Finetuning of Quantized LLMs
A 2023 breakthrough that combined 4-bit quantization with LoRA, making it possible to fine-tune massive 65-billion parameter models on a single consumer GPU.
Read breakdownReAct
ReAct: Synergizing Reasoning and Acting in Language Models
A 2022 framework that interleaves reasoning (Chain-of-Thought) with acting (using external tools like Wikipedia or APIs), forming the basis of modern LLM agents.
Read breakdownRetrieval-Augmented Generation (RAG)
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
The 2020 paper that introduced a general-purpose fine-tuning recipe for combining pre-trained parametric and non-parametric memory.
Read breakdownRoFormer
RoFormer: Enhanced Transformer with Rotary Position Embedding
Introduced Rotary Position Embedding (RoPE), a method that mathematically integrates absolute positional information with relative distances, becoming the standard for modern LLMs.
Read breakdownScaling Laws for Neural Language Models
Scaling Laws for Neural Language Models
Demonstrated that language model loss decreases predictably as a power law with compute, dataset size, and parameter count, establishing the mathematical foundation for the LLM scaling race.
Read breakdownScaling Monosemanticity
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Anthropic's 2024 paper successfully applied Sparse Autoencoders to Claude 3 Sonnet, extracting millions of high-level, interpretable concepts (like 'Golden Gate Bridge' or 'security vulnerabilities') from a state-of-the-art production model.
Read breakdownScore-Based SDEs
Score-Based Generative Modeling through Stochastic Differential Equations
The 2020 paper that unified diffusion models and score-matching models into a single, elegant framework based on continuous-time Stochastic Differential Equations (SDEs).
Read breakdownSegment Anything Model (SAM)
Segment Anything
The 2023 Meta paper that introduced a promptable foundation model for image segmentation, capable of zero-shot segmentation of any object.
Read breakdownSelf-Consistency
Self-Consistency Improves Chain of Thought Reasoning in Language Models
A 2022 paper that improved upon Chain-of-Thought prompting by generating multiple reasoning paths and taking a majority vote on the final answer.
Read breakdownSelf-Instruct
Self-Instruct: Aligning Language Models with Self-Generated Instructions
The 2022 paper that demonstrated how to use a large, proprietary LLM to generate synthetic training data to fine-tune smaller, open-source models.
Read breakdownSentence-BERT (SBERT)
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
The 2019 paper that adapted BERT for generating semantically meaningful sentence embeddings that can be compared using cosine similarity.
Read breakdownSigLIP
Sigmoid Loss for Language Image Pre-Training
The 2023 paper that replaced CLIP's softmax contrastive loss with a simpler sigmoid loss, allowing for massive scaling and better performance.
Read breakdownSpider
Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task
The 2018 dataset paper that defined the cross-domain text-to-SQL task and became the standard benchmark for natural language database querying.
Read breakdownSwin Transformer
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
The 2021 paper that brought hierarchical structure and local windows to Vision Transformers, making them practical for high-resolution tasks like object detection and segmentation.
Read breakdownSwitch Transformer
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Scaled a sparse Mixture of Experts (MoE) model to a trillion parameters, proving that massive parameter scaling is possible without a proportional increase in compute.
Read breakdownT5
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Introduced T5, a framework that reframes every NLP task into a text-to-text format, enabling a single model architecture to handle diverse tasks.
Read breakdownToolformer
Toolformer: Language Models Can Teach Themselves to Use Tools
A 2023 Meta paper that taught models to use external tools by automatically generating their own training data, fine-tuning the model to natively call APIs via special text tokens.
Read breakdownTowards Monosemanticity
Towards Monosemanticity: Decomposing Neural Networks with Sparse Autoencoders
A 2023 mechanistic interpretability paper by Anthropic that successfully used Sparse Autoencoders to extract understandable, human-readable concepts from the impenetrable 'black box' of neural network activations.
Read breakdownTrain Short, Test Long
Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
Introduced ALiBi (Attention with Linear Biases), a simple positional encoding scheme that allowed models to extrapolate to context lengths far longer than what they were trained on.
Read breakdownTraining Compute-Optimal LLMs
Training Compute-Optimal Large Language Models
The 'Chinchilla paper' that proved models should be scaled equally with training data, overturning the previous consensus that large models could be trained on relatively little data.
Read breakdownTraining LMs to Follow Instructions with Human Feedback
Training language models to follow instructions with human feedback
The 2022 paper from OpenAI that popularized RLHF (Reinforcement Learning from Human Feedback), creating models like InstructGPT and paving the way for ChatGPT.
Read breakdownTree of Thoughts
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
A 2023 paper that framed LLM reasoning as a search algorithm over a tree of possible thoughts, allowing models to evaluate their own progress, backtrack, and look ahead.
Read breakdownVision Transformer (ViT)
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
The 2020 paper that proved Transformer architectures could replace CNNs for computer vision by treating image patches as a sequence of words.
Read breakdownVQ-VAE
Neural Discrete Representation Learning
The 2017 DeepMind paper that introduced the Vector Quantized Variational Autoencoder, turning continuous images into sequences of discrete, learnable 'tokens'.
Read breakdownwav2vec 2.0
wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
The 2020 Facebook AI paper that successfully applied self-supervised learning to audio, allowing speech recognition models to be trained with vastly less human-transcribed data.
Read breakdownWhisper
Robust Speech Recognition via Weak Supervision
The 2022 OpenAI paper that achieved human-level robustness in speech recognition by training on a massive 680,000-hour dataset of noisy, weakly supervised web audio.
Read breakdownZeRO
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Introduced the Zero Redundancy Optimizer, a memory optimization technology that partitions model states across GPUs, making it possible to train models with hundreds of billions of parameters.
Read breakdownDeep Residual Learning
Deep Residual Learning for Image Recognition
The 2015 paper that introduced residual skip connections, solving the vanishing gradient problem and making extremely deep networks trainable.
Read breakdownImageNet Classification with Deep CNNs
ImageNet Classification with Deep Convolutional Neural Networks
The 2012 paper (AlexNet) that proved deep convolutional neural networks trained on GPUs could crush traditional computer vision methods.
Read breakdownBatch Normalization
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Introduced Batch Normalization, a technique that normalized activations across a mini-batch, making networks faster and dramatically more stable to train.
Read breakdownDropout
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Introduced Dropout, a remarkably simple yet highly effective regularization technique that prevents neural networks from overfitting by randomly disabling neurons.
Read breakdownAdam Optimizer
Adam: A Method for Stochastic Optimization
Introduced Adam, an adaptive optimization algorithm that combined the best properties of momentum and RMSProp, becoming the default optimizer for deep learning.
Read breakdownLayer Normalization
Layer Normalization
Introduced Layer Normalization, which normalizes activations across the features of a single data point rather than across a batch, making it ideal for sequence models.
Read breakdownSequence to Sequence Learning
Sequence to Sequence Learning with Neural Networks
Introduced the Seq2Seq encoder-decoder architecture, allowing neural networks to map input sequences to output sequences of entirely different lengths.
Read breakdownNeural Machine Translation (Attention)
Neural Machine Translation by Jointly Learning to Align and Translate
Introduced the Attention mechanism, allowing Sequence-to-Sequence models to dynamically 'look back' at relevant parts of the input sentence instead of relying on a single context vector.
Read breakdownU-Net
U-Net: Convolutional Networks for Biomedical Image Segmentation
Introduced U-Net, an elegant, symmetric architecture for image segmentation that remains the foundational backbone for modern diffusion models.
Read breakdownWord2Vec
Efficient Estimation of Word Representations in Vector Space
Introduced Word2Vec, a highly efficient method for learning dense word embeddings that captured semantic meaning and analogical relationships.
Read breakdownRandom Forests
Random Forests
Formalized Random Forests, an ensemble learning method that combines bagging and random feature selection to build highly robust and accurate predictive models.
Read breakdownXGBoost
XGBoost: A Scalable Tree Boosting System
Introduced XGBoost, a highly scalable and regularized gradient boosting library that utterly dominated Kaggle competitions and tabular data modeling.
Read breakdownGenerative Adversarial Networks (GAN)
Generative Adversarial Nets
Introduced the GAN, an elegant framework where two neural networks—a generator and a discriminator—compete against each other to create hyper-realistic synthetic data.
Read breakdownVariational Autoencoders (VAE)
Auto-Encoding Variational Bayes
Introduced the Variational Autoencoder, a generative model that learns a continuous, structured latent space through the mathematical 'reparameterization trick'.
Read breakdownProximal Policy Optimization (PPO)
Proximal Policy Optimization Algorithms
Introduced PPO, an incredibly stable and sample-efficient Reinforcement Learning algorithm that became the default standard, eventually powering RLHF in ChatGPT.
Read breakdownSHAP (SHapley Additive exPlanations)
A Unified Approach to Interpreting Model Predictions
Introduced SHAP, a unified framework for interpreting complex machine learning models by applying game theory to calculate the exact marginal contribution of each feature.
Read breakdownKnowledge Distillation
Distilling the Knowledge in a Neural Network
Formalized Knowledge Distillation, a method for training small, fast 'student' models to mimic the nuanced behavior of massive 'teacher' models.
Read breakdown