GitHub - mizorewww/laya-mlx: Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
Open-weight typed decisions, running natively on Apple Silicon. 13.4 ms median end-to-end for a short English typed decision. 7.4 ms with the multilingual checkpoint. 0 output tokens. Local MLX inference, with no PyTorch, Transformers runtime, or cloud API. 中文 · Benchmarks · Snake demo · Huggin...
はてなテクノロジー
2026年09月20日 09:11