Transformers
Tag2.1K stories
Transformers news and research on the architecture behind current language and vision models. Readers can learn about attention mechanisms and positional encoding, pretraining and fine-tuning, efficiency variants, and applications across text and multimodal tasks.
Understanding Long RoPE in LLMsDecoding Complexity with Transformers: Researchers from Anthropic Propose a Novel Mathematical Framework for Simplifying Transformer ModelsThe Art of Memory Mosaics: Unraveling AI’s Compositional ProwessFine-Tuning AI Models for Extractive Question Answering in ElixirHow ‘Chain of Thought’ Makes Transformers Smarter[2405.00738] HLSTransform: Energy-Efficient Llama 2 Inference on FPGAs Via High Level SynthesisData Science Unicorns, RAG Pipelines, a New Coefficient of Correlation, and Other April Must-ReadsSelf-Attention in Transformers: Computation Logic and ImplementationFine-tuning Llama-3 with ORPO: A Deep DiveHands-on: this self-transforming Megatron is as badass as it is expensive