AI Inference
Tag960 stories
AI Inference news and updates covering the stage where a trained model serves predictions, as distinct from training. Readers can learn about serving frameworks, batching and KV caching, quantization, latency and throughput tuning, accelerator selection, and the cost of running models in production.
Serving LLMs on an RTX4090 with SequoiaA Hitchhiker’s Guide to Speculative DecodingHuawei AI Introduces ‘Kangaroo’: A Novel Self-Speculative Decoding Framework Tailored for Accelerating the Inference of Large Language ModelsLayerSkip: An End-to-End AI Solution to Speed-Up Inference of Large Language Models (LLMs)Powerful ASR + diarization + speculative decoding with Hugging Face Inference EndpointsTurbocharging Meta Llama 3 Performance with NVIDIA TensorRT-LLM and NVIDIA Triton Inference ServerResearchers at CMU Introduce TriForce: A Hierarchical Speculative Decoding AI System that is Scalable to Long Sequence GenerationEffort EngineResearchers at Apple Propose ReDrafter: Changing Large Language Model Efficiency with Speculative Decoding and Recurrent Neural NetworksNVIDIA TensorRT Accelerates Stable Diffusion Nearly 2x Faster with 8-bit Post-Training Quantization