2 September 2026
Apple releases engine optimizing AI models on its chips
First reported
TLDR AI ran this on .
- Apple built Lily, software that runs large language models directly on Apple devices rather than sending data to remote servers.
- Lily uses Apple's unified memory architecture, which lets the processor and graphics chip share data more efficiently than typical setups.
- Testing showed Lily processes text faster than MLX-LM, an existing open-source tool, when running Qwen 3.6-35B, a Chinese AI model with 35 billion parameters.
How it was covered
TLDR AITLDR editorial team
Apple's Lily engine optimizes on-device LLM inference for Apple silicon by leveraging unified memory and specialized hardware, outperforming MLX-LM in prefill and decode throughput for models like Qwen3.6-35B.