Cross-Family Speculative Decoding Boosts Polish LLM Performance on Apple Silicon
Researchers extended the MLX-LM framework to enable cross-tokenizer speculative decoding, improving inference speed for Polish language models on Apple Silicon. The study evaluated Bielik-11B-Instruct paired with three draft models, showing significant speedups.