Researchers at Apple have proposed a method called ReDrafter, which combines speculative decoding with recurrent neural networks to improve the efficiency of large language models (LLMs) in generating text. ReDrafter allows for faster response generation without compromising the model's depth or output quality.