local-first inference

A model, running on your machine.

This is Tier 1 of Aperture's routing — inference that happens entirely on-device via WebGPU. Weights download once and cache in your browser; after that, nothing you type leaves your computer.

Pick a model and run it entirely on your device — no data leaves the browser. Weights download once (0.9 GB) and cache.

requires a WebGPU browser (recent chrome, edge, or safari) · pick a model below; the first load downloads its weights and caches them