2 September 2026

Apple releases engine optimizing AI models on its chips

First reported

TLDR AI ran this on .

  • Apple built Lily, software that runs large language models directly on Apple devices rather than sending data to remote servers.
  • Lily uses Apple's unified memory architecture, which lets the processor and graphics chip share data more efficiently than typical setups.
  • Testing showed Lily processes text faster than MLX-LM, an existing open-source tool, when running Qwen 3.6-35B, a Chinese AI model with 35 billion parameters.

How it was covered

TLDR AITLDR editorial team

Apple's Lily engine optimizes on-device LLM inference for Apple silicon by leveraging unified memory and specialized hardware, outperforming MLX-LM in prefill and decode throughput for models like Qwen3.6-35B.