LiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices.
-
Updated
Aug 23, 2026 - C++
LiteRT-LM is Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices.
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
[Official] prima.cpp: Fast 30-70B LLM inference on heterogeneous and everyday home devices
Open-source production-ready mobile server-less AI apps. Clone, customize, and ship on Android & iOS with MLange.
On-device shell command generator for macOS Tahoe. Uses Apple's 3B model with dynamic few-shot retrieval from 21k tldr examples.
KnoLo Core is a local-first knowledge base engine built for small language models (LLMs). It packages your documents into a compact .knolo file and enables fully deterministic querying — no embeddings, no vector databases, no cloud services required. Designed for on-device and edge LLM deployments.
DeepSeek-V4-Flash-0731 284B inference in ~30 GB of RAM on any M-series MacBook
Agentic Android Open Source Project (AAOSP) — Android fork with native LLM system service, MCP-aware apps, and an agent-driven launcher. On-device Qwen 2.5 via llama.cpp. Apps declare tools in their manifest. The OS runs the model.
Declare the outcome, skip the logic. The developer-first infrastructure engine for structured, multi-turn AI conversations with built-in safety, async follow-ups, and native compliance
Android 16 fork. AI as a platform primitive. Twelve capabilities, one shared runtime, every app. OEM-pluggable. Apache 2.0.
Rivo Agent is a private, local-first mobile AI assistant that runs compact GGUF language models directly on your phone. It lets users download a compatible model, chat offline, keep local memory, customize assistant behavior, and use AI privately without sending conversations to a remote inference server.
Apple FoundationModels API on iOS 18+. Same call site, native passthrough on iOS 26 (Apple Intelligence), CoreML / MLX backends on older OSes. Drop-in source compatible.
Curated resource for mobile teams shipping on-device LLMs — runtime benchmarks, model picks, GDPR-friendly architecture, and real production use cases.
Run LLMs on Snapdragon NPU — including the 'unsupported' 8 Gen 1 (Hexagon v69). Verified at 31 tok/s on OnePlus 10 Pro.
Kotlin Multiplatform engine for running Gemma LLMs on-device on Android via LiteRT-LM — stateful KV-cache chat sessions, resumable model management, function calling. Includes NativeLM, a private Local AI chat app. AGPL-3.0 / commercial.
OnDevice Local AI Studio — run LLMs on iPhone offline. SwiftUI workbench with MLX, llama.cpp, whisper.cpp, RAG, and authenticated OpenAI/Anthropic/Ollama-compatible local APIs.
High-performance Android SDK for on-device LLM inference (GGUF). Privacy-focused, offline-first, and powered by llama.cpp with a clean Kotlin Coroutines API.
Reverse-engineering notes on fm, Apple's Foundation Models CLI in macOS 27: on-device model catalog (9M/85M/300M/3B + code/vision/speech), Private Cloud Compute, Siri local<->cloud routing, and the OpenAI-compatible 'fm serve' API.
📱 手机端 AI 操作系统全景知识库 — 334+ 篇深度页面,覆盖端侧大模型、AI Agent、芯片适配、推理优化 | 自动更新
Add a description, image, and links to the on-device-llm topic page so that developers can more easily learn about it.
To associate your repository with the on-device-llm topic, visit your repo's landing page and select "manage topics."