Small Models, Big Shift: Why On-Device AI Is Now Rivalling the Cloud
Compact AI models running directly on phones are proving that capable intelligence doesn't always need a data centre, bringing gains in privacy, speed, and energy efficiency.
For years, the AI industry has worked on the assumption that bigger is better: more parameters, more GPUs, more data centres. That thinking is now being tested. A new generation of compact models that run entirely on smartphones is showing that useful, dependable AI can live in your pocket, with no cloud connection required.
The Case Against "Bigger Is Better"
A recent UNESCO report urged governments and companies to move away from energy-hungry AI and invest in leaner, more sustainable models, estimating that such changes could cut energy use by as much as 90%. While much of the industry keeps pouring money into new data centres, on-device AI offers a very different path.
Several forces are pushing this shift. Users increasingly want their data to stay private, concerns about the environmental cost of AI keep growing, and many real-world tasks simply need to run faster and more efficiently than a round trip to the cloud allows. Two technical advances have made it practical: modern smartphone chips now include dedicated AI accelerators, and improved training methods allow much smaller models to deliver quality that once demanded billions of extra parameters.
A New Compact Vision Model Enters the Race
Tether AI Research has open-sourced VisionPsy-Nano, a small vision-language model designed specifically for phones and edge devices. According to the company, it beats similar-sized rivals, including Liquid AI's LFM2.5-VL-450M and Hugging Face's SmolVLM2-500M.
The bigger point is what running locally makes possible. Because everything happens on the device, photos and questions never leave the phone. There is no upload, no server-side log, and the feature keeps working even when there is no internet connection. Privacy is built into the design rather than promised in a policy.
What a Vision-Language Model Actually Does
A vision-language model (VLM) combines an image encoder with a text-based language model, so it can look at visual information and reason about it in words. Even the biggest cloud players have moved into this space with lighter models, with Google's open-weight family progressing from PaliGemma in 2024 to the Gemma 3 and Gemma 4 lines.
For smartphones, though, the real action is at roughly half a billion parameters. Models at this size can fit inside a phone's memory and power limits while still performing reliably. It has become one of the busiest areas in open model development, with Liquid AI's LFM vision models and Hugging Face's SmolVLM and nanoVLM projects among the closest competitors, and new architectures and training methods arriving quickly.
Why Latency Matters So Much
Latency is the gap between asking an AI system something and getting an answer, whether through a chatbot or a voice assistant. The shorter that gap, the better the experience. For casual use, a lag is merely irritating. In time-critical settings such as medical diagnosis, emergency response, or navigation, a slow answer can cause real harm, and in the worst cases it can lead to serious operational failures.
The next wave of AI may not be defined by who owns the biggest data centre, but by who can fit genuine intelligence into the devices people already carry.
The Road to Embodied AI and Robotics
Beyond everyday tasks like reading a document, interpreting a chart, or describing a scene, small VLMs are being studied for embodied AI, where machines must perceive and act in the physical world. Real-time augmented reality, robotics, and brain-computer interfaces all need near-instant responses.
A key example is scene visual question answering, where a system continually assesses its surroundings to answer spatial questions. Developers building these edge systems accept a small drop in general reasoning accuracy in return for the very low latency and efficiency needed to run locally. The payoff is machines that can see, think, and act in real time without relying on a cloud connection. VisionPsy-Nano fits into a wider goal of making AI accessible to everyone by running it fully on device.
- Bigger is no longer the only path. Compact models can deliver capable AI without large cloud infrastructure, and could cut energy use by up to 90%.
- Hardware and training have caught up. Built-in AI accelerators on phones and better training methods let small models punch above their weight.
- On-device means private by design. Data never leaves the phone, and features keep working offline.
- Speed can come with little quality loss. The Flash variant reports up to 23x faster first responses while keeping about 99% of quality.
- Embodied AI is the next frontier. Robotics, AR, and brain-computer interfaces need the low latency that local models can offer.
