Perplexity open-sources Lily inference engine for Apple silicon
Lily is specialized for Qwen3.6-35B-A3B to handle local inference for hybrid compute in Perplexity Computer.
- Lily separates prefill and decode into distinct workloads to optimize memory bandwidth and weight reuse on Apple silicon.
- Benchmarks on an M5 Max MacBook Pro showed an average 1.23× higher prefill throughput and 1.35× higher decode throughput over MLX-LM.
- The code is available in the perplexityai/pplx-garden GitHub repository.
Developers running Qwen models locally on Mac hardware can achieve higher throughput without sacrificing output quality.

Sources
Read this as text
Back to the AI news