Efficient conversion, runtime, and optimization for on-device machine learning.
LiteRT isn't just new; it's the next generation of the world's most widely deployed machine learning runtime. It powers the apps you use every day, delivering low latency and high privacy on billions of devices.

Trusted by the most critical Google apps

100K+ applications, billions of global users

LiteRT Highlights

Deploy via LiteRT

Streamline your deep learning workflow from training to on-device deployment.
Use .tflite pre-trained models or convert PyTorch, JAX or TensorFlow models to .tflite.
Use the LiteRT optimization toolkit to quantize your models post-training.
Deploy your model with LiteRT and pick the optimal accelerator for your app.

Choose Your Development Path

Use LiteRT to deploy AI anywhere—from high-performance mobile apps to resource-constrained IoT devices.
Transitioning to LiteRT to leverage enhanced performance and unified APIs across platforms (Android, Desktop, Web).
Have a PyTorch model, looking to implement on-device vision or audio experiences.
Creating sophisticated on-device chatbots using optimized open-weight GenAI models like Gemma or another open-weight model with LiteRT-LM.
Authoring custom models or performing deep hardware-specific CPU/GPU/NPU optimizations for peak performance.

Samples, models, and demo

Complete, end-to-end sample apps.
Pre-trained, out-of-the-box Gen AI models.
A gallery that showcases on-device ML/GenAI use cases using LiteRT.

Blogs and Announcements

Stay up to date with the latest announcements, technical deep dives, and performance benchmarks from the LiteRT team.
LiteRT.js, Google's high-performance web AI inference library for running models using WebAssembly, WebGPU, and WebNN.