Ollama
Run open LLMs locally with a single command.
Small specialised on-device models for speech, text, and vision, one per task
This is infrastructure for app developers rather than a tool you sit down and use, and on those terms it is well judged. Small per-task models running on-device remove the three costs that make cloud inference painful in consumer apps: latency, per-call billing, and the privacy conversation. The 100,000 monthly active device threshold is generous enough that most apps will never pay. Two things to be precise about. The licence is source-available, explicitly not open source, and attribution is required, so read it before shipping. And commercial pricing above the threshold is not published, which makes the cost of success unknowable in advance.
Desert Ant Labs publishes small task-specific models that run entirely on-device across speech, text, and vision, delivered via Swift, Kotlin, and JavaScript SDKs and free below 100,000 monthly active devices.
Desert Ant Labs takes the opposite bet to frontier model vendors: rather than one large model doing everything through a cloud API, it publishes many small models that each do one job and run on the user device. The catalogue includes Voz for speech recognition, Align, Clear, Uhm, Ear, and Clips for audio, Redact, Gist, Title, Tongue, and Emo for text, and Shapes for vision, with further models in beta. The practical consequence is that there is no cloud API and no per-call cost. You embed the model via a Swift, Kotlin, or JavaScript SDK and it executes locally, which removes latency, works offline, and means user audio or text never leaves the handset. Weights are published on Hugging Face. Licensing is source-available rather than open source. The Desert Ant Labs Source-Available Licence 1.0 permits free use below 100,000 monthly active devices per platform and per model, with no limit on how often each person runs it and no per-call fee; research and teaching are exempt from the threshold. Above it, a commercial licence is required. The company is a Dutch entity, Desert Ant Labs B.V.
Desert Ant Labs publishes small, one-task-each models that run on-device, free below 100,000 monthly active devices.
Desert Ant Labs B.V. is a Dutch company publishing small specialised models designed to run on end-user devices rather than in the cloud.
It was featured on Product Hunt in September 2026 and distributes model weights publicly on Hugging Face under a source-available licence.
The catalogue spans audio models including Voz, Align, Clear, Uhm, Ear, and Clips, text models including Redact, Gist, Title, Tongue, and Emo, and vision models including Shapes, with further beta models.
Delivery is via Swift, Kotlin, and JavaScript SDKs with no cloud API and no per-call cost. Use is free below 100,000 monthly active devices per platform and model.
Mobile and web application developers who need AI capability without cloud latency, per-call billing, or sending user data off-device.
Developers embedding models in applications.
Product and engineering leads at app companies.
Edge AI and on-device machine learning communities.
A consumer app team that needs speech or text processing at scale without per-call inference costs or a privacy review.
Not stated; the entity is Desert Ant Labs B.V., governed by Netherlands law.
No. They run entirely on the device via Swift, Kotlin, or JavaScript SDKs, with no cloud API.
Above 100,000 monthly active devices per platform and per model, which requires a commercial licence from the vendor.
No. It uses the Desert Ant Labs Source-Available Licence 1.0, which the vendor explicitly describes as source-available rather than open source.
Speech recognition, alignment, noise clearing, filler removal, clipping, redaction, summarisation, titling, translation, emotion, and vision, with more models in beta.
Model weights are published publicly on Hugging Face.
Side-by-side pages for pricing, features, and best-fit use cases.
Run open LLMs locally with a single command.
One API for hundreds of AI models across providers.
AI cloud with 200+ model APIs, serverless inference and GPU instances
Serverless GPU runtime for AI inference, training, and sandboxes