2 links tagged with all of: transformer + tokenization + attention
Click any tag below to further narrow down your results
Links
This article traces the path from RNNs to transformers, explaining why attention-based, non-recurrent architectures replaced older models. It then breaks down encoder vs decoder designs and shows how GPT’s decoder-only approach and tokenization power today’s large language models.
- Transformers replaced RNNs by dropping recurrence entirely, using self-attention to compute all token relationships in parallel and eliminating both the sequential bottleneck and vanishing/exploding gradient issues.
- The original transformer architecture split into two lineages: encoder-only (BERT) for contextual understanding, used in Google Search ranking, and decoder-only (GPT) for text generation.
- GPT's scale jumped dramatically across versions—GPT-2 had 10x GPT-1's parameters/data, GPT-3 scaled another 100x—and GPT-3 showed strong zero-shot/few-shot performance without needing task-specific fine-tuning.
This article breaks down how an LLM turns your prompt into streamed tokens, covering tokenization, embeddings, transformer attention, and the two-phase pipeline of compute-bound prefill and memory-bound decode. It explains KV caching, quantization, and metrics like Time to First Token and Inter-Token Latency to show why inference speed depends on both compute and memory.
- LLM inference splits into compute-bound prefill (TTFT) and memory-bound decode (ITL), which explains why the two phases have fundamentally different performance bottlenecks.
- KV caching avoids quadratic recomputation but costs real memory—about 1 MB per token for a 13B model, so a 4,000-token context burns 4 GB of VRAM.
- Techniques like INT8/INT4 cache quantization, token eviction, and paged cache management exist specifically to tame this KV cache memory growth.
- DeepSeek's V4 cuts KV cache size by 90% at a million-token context by combining sparse and dense compressed attention.