Skip to content
Back to skills

Shipping A Model In An Android App

ASecurity

Put a local language model inside a native Android app written in Kotlin or Java. Covers MediaPipe LLM Inference and LiteRT-LM for .task and .litertlm models, llama.cpp through the NDK for GGUF, ML Kit GenAI and Gemini Nano where a built-in model is enough, GPU and NNAPI delegation, Play asset packs, and whether to bundle the model or download it on first run. Use when the project is a native Android app and someone wants on-device or offline AI or local inference. Not for React Native or Exp...

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 19, 2026
ai-agentsgojavaswiftkotlinreactdebuggingapibackenddocumentation

Works with

  • cli
  • api

Security analysis

A100/100

Pro scans all 4 files and shows the line behind each finding

Scanned September 19, 2026

npx -y skills add ErtasAI/open-model-skills --skill shipping-a-model-in-an-android-app --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Shipping A Model In An Android App?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Shipping A Model In An Android App
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/ertasai-shipping-a-model-in-an-android-app/badge)](https://www.skillsdirectory.com/skills/ertasai-shipping-a-model-in-an-android-app)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: shipping-a-model-in-an-android-app
description: >-
  Put a local language model inside a native Android app written in Kotlin or
  Java. Covers MediaPipe LLM Inference and LiteRT-LM for .task and .litertlm
  models, llama.cpp through the NDK for GGUF, ML Kit GenAI and Gemini Nano where
  a built-in model is enough, GPU and NNAPI delegation, Play asset packs, and
  whether to bundle the model or download it on first run. Use when the project
  is a native Android app and someone wants on-device or offline AI or local
  inference. Not for React Native or Expo projects, not for iOS native Swift
  projects, not for Flutter, and not for models running on a server.
license: Apache-2.0
metadata:
  version: "0.1.0"
  author: "Edward Xi Yang, Ertas AI"
  stage: "ship"
  previous_skill: "inspecting-a-model-bundle"
---

# Shipping a model in an Android app

Android splits into two runtime families that take different artifact
shapes, plus a third option that ships no model at all. Which of the two
runtimes fits depends on what file is already in hand; which of the three
paths fits depends on whether a built-in model is actually good enough for
the task. Settle both before installing anything.

## Which artifact shape this needs

| You are holding | Package | What it eats |
|---|---|---|
| A `.task` or `.litertlm` file | MediaPipe LLM Inference, or LiteRT-LM's own Android API | A Task Bundle or a LiteRT LM file, produced by `litert-torch` |
| A single `.gguf` file | llama.cpp through the NDK | GGUF, self-contained |

These are two separate native stacks with two separate build paths. A
`.litertlm` file does not load into a llama.cpp-based wrapper, and a `.gguf`
file does not load into MediaPipe's `LlmInference` or LiteRT-LM's `Engine`.

**If what you are holding is neither of these,** a merged Hugging Face
checkpoint (`config.json` + `model*.safetensors`) or a PEFT adapter directory
(`adapter_config.json` + `adapter_model.safetensors`), it is not shippable
into a native Android app as is. Run **inspecting-a-model-bundle** first to
confirm which shape you actually have, then convert:

- To `.litertlm`: `litert-torch export_hf`, covered below and in
  `references/mediapipe-litert-path.md`.
- To GGUF: `convert_hf_to_gguf.py` then `llama-quantize`, covered below and
  in `references/llamacpp-jni-path.md`.

A PEFT adapter directory converts to neither format directly. Merge it into
its base model first with `peft`'s `merge_and_unload()`, then export from the
merged checkpoint. The GGUF path has a second option: convert the adapter
itself with `convert_lora_to_gguf.py` and load it as a runtime LoRA against a
base GGUF through the raw llama.cpp C API, though see the platform gotchas
below for how much of that API the common Android wrappers actually expose.

**If it is genuinely unclear which of the two to pick,** the deciding factor
is usually whether the model family is already covered by Google's
pre-converted set. MediaPipe LLM Inference ships pre-converted Gemma-3n,
Gemma-3 1B, Gemma-2 2B, and Phi-2 models, which makes it the fastest path for
those families and for anything else exported through `litert-torch`.
llama.cpp through the NDK is the better default otherwise: GGUF is the
highest-coverage artifact across the whole open-model ecosystem, quantisation
tooling is mature, and nothing about the model family limits which base
models it accepts.

## Install and wiring

Every quickstart below is shown in Kotlin, matching what Google's own docs
and llama.cpp's official Android example both use. All three are plain
Android or JVM libraries, so a Java call site is mechanically possible, but
no Java-specific sample was verified for any of the three paths in this
pass; translate the builder chains and lambdas by hand if the project is
Java-only.

### MediaPipe LLM Inference (.task, .litertlm)

```gradle
implementation 'com.google.mediapipe:tasks-genai:0.10.35'
```

Kotlin, from Google's own LLM Inference guide, loading a model pushed to the
device with `adb`:

```kotlin
val taskOptions = LlmInferenceOptions.builder()
        .setModelPath("/data/local/tmp/llm/model_version.task")
        .setMaxTopK(64)
        .build()

val llmInference = LlmInference.createFromOptions(context, taskOptions)
```

`setModelPath` also accepts a `.litertlm` file; the same `LlmInference`
class loads either.

### LiteRT-LM's own Android API (.litertlm)

```gradle
implementation("com.google.ai.edge.litertlm:litertlm-android:0.14.0")
```

LiteRT-LM ships a separate, native Kotlin API distinct from MediaPipe's
`tasks-genai`, documented as the recommended entry point for native Android
apps and JVM-based desktop tools. Loading and streaming, from Google's own
LiteRT-LM Android documentation:

```kotlin
val engineConfig = EngineConfig(modelPath = "/path/to/model.litertlm")
val engine = Engine(engineConfig)
engine.initialize()
```

### llama.cpp through the NDK (GGUF)

Unlike React Native's `llama.rn` or iOS's `LLM.swift`, there is no single
mainstream Maven-published wrapper that most native Android GGUF projects
converge on. llama.cpp's own repository carries an official Android example
project at `examples/llama.android`, meant to be built as part of a full
llama.cpp source checkout rather than pulled in as a Gradle dependency: its
CMake configuration adds the llama.cpp source tree as a subdirectory several
levels up from the example project itself. The practical route is to clone
llama.cpp and use that example as the template for a native Android module,
or fork its JNI layer into an existing app.

If pulling in a prebuilt dependency instead of building from source is
preferred, `Llamatik` is a community Kotlin Multiplatform wrapper around
llama.cpp published to Maven Central that does not require any NDK setup of
its own:

```kotlin
commonMain.dependencies {
    implementation("com.llamatik:library:1.9.1")
}
```

1.9.1 is the latest version on Maven Central, checked 2026-07-29; the
project's own README still shows 1.7.0 in its install snippet, so check
Maven Central rather than the README for the current coordinate. Minimum
Android API level 26. This is a third-party project, not part of the
llama.cpp repository itself; verify it still meets the project's needs before
depending on it, the way any third-party dependency should be evaluated.

## A minimal working example

### MediaPipe LLM Inference: load, generate, stream

Synchronous, from Google's own guide:

```kotlin
val result = llmInference.generateResponse(inputPrompt)
```

Streaming, registering a result listener at options-build time and then
calling the async variant:

```kotlin
val options = LlmInference.LlmInferenceOptions.builder()
    .setModelPath("/data/local/tmp/llm/model_version.task")
    .setResultListener { partialResult, done ->
        // partialResult is the newly generated text fragment;
        // done is true on the final callback for this request
    }
    .build()

val llmInference = LlmInference.createFromOptions(context, options)
llmInference.generateResponseAsync(inputPrompt)
```

`generateResponseAsync` returns immediately; the result listener fires once
per streamed chunk, with `done = true` marking the end of that generation.

### LiteRT-LM: load, generate, stream

From Google's own LiteRT-LM Android documentation, using a conversation
object and Kotlin's coroutine `Flow`, described as the recommended approach
for coroutine users:

```kotlin
engine.createConversation().use { conversation ->
    conversation.sendMessageAsync("What is the meaning of life?")
        .collect { chunk ->
            // chunk is the newly generated piece; append it to UI state here
        }
}
engine.close()
```

`sendMessageAsync` returns a `Flow` that emits as generation proceeds;
`collect` is where a UI would append each piece to the growing response.

### llama.cpp through the NDK: load, generate, stream

From llama.cpp's own official Android example, whose Kotlin wrapper exposes
a suspend-based load call and a `Flow`-based streaming call, mirroring the
LiteRT-LM shape above despite being a completely different native stack
underneath:

```kotlin
val engine = AiChat.getInferenceEngine(applicationContext)
engine.loadModel(modelFile.path)
engine.setSystemPrompt("You are a helpful assistant.")

engine.sendUserPrompt("What is the meaning of life?")
    .onCompletion { /* generation finished */ }
    .collect { token ->
        // token is the newly generated piece; append it to UI state here
    }

engine.destroy()
```

`loadModel` and `setSystemPrompt` are `suspend` functions; `sendUserPrompt`
takes an optional `predictLength` parameter (default 1024 tokens) and
returns a `Flow<String>` that emits one token at a time. Call `destroy()`
when the screen or activity holding the engine is torn down to release the
native model.

## Or ship nothing at all

Before installing any of the above, check whether the task fits inside
what's already on the device. ML Kit's GenAI APIs sit on top of Gemini Nano,
Google's own on-device foundation model, reached through Android's AICore
system service rather than a model you bundle or host. Google's own
description: ML Kit's GenAI APIs "harness the power of Gemini Nano to help
your apps perform tasks," built on AICore, "an Android system service that
enables on-device execution of GenAI foundation models."

The APIs on offer are narrow and task-specific, not a general chat endpoint:
Prompt (free-text and multimodal prompting), Summarization, Proofreading,
Rewriting, Image Description, and Speech Recognition. For a task that fits
one of these cleanly, this whole skill is unnecessary: no model to bundle or
download, no size budget to plan around, and the model stays current as
Google updates it.

Device eligibility is the real constraint, and the figures below come from
reports summarising Google's device-support statements rather than a single
authoritative page, so treat them as indicative: Gemini Nano availability is
commonly reported at over 140 million devices, covering
the Pixel 8 and 9 series, Galaxy S24 and S25 series, Z Fold and Flip 6, and
recent Motorola Razr models. Treat this list as broadly right rather than
exact, and check the current device list before promising a feature works on
a specific phone. On an ineligible device the ML Kit GenAI APIs are reachable
but unavailable, so a real app still needs a fallback: either a degraded
feature or one of the shipped-model routes above.

## The delivery decision

Once the model runs, the next question is how it reaches the device: bundled
into the app, delivered through a Play asset pack, or downloaded on first
run. Full detail, the app store size ceilings for both platforms, and the
first-run download experience are in
`references/model-delivery-and-size-budgets.md`.

**This is where Android and iOS genuinely diverge, not just in numbers but
in what's viable at all.** iOS has one practical delivery mechanism beyond
the base app binary: download on first run, bounded by whatever App Store
review and install-conversion tolerance allow. Android has a second,
first-class mechanism that iOS has no equivalent of: Play asset packs,
delivered install-time, fast-follow, or on-demand, with ceilings that sit far
above what a single iOS binary can carry; the exact figures for each
delivery mode are in `references/model-delivery-and-size-budgets.md` rather
than repeated here. Bundling a model at a size that would be unthinkable on
iOS is a normal, supported path on Android through an on-demand asset pack,
requested at runtime rather than downloaded through app-side networking
code. From Google's own Play Asset Delivery guide, fetching a pack and
tracking its progress, assuming an already-obtained `assetPackManager`
(that page does not show how the instance itself is constructed, so confirm
that specific call against the Play Core library reference before writing
it):

```kotlin
assetPackManager.requestFetch(listOf("model_pack"))
assetPackManager.registerListener { state ->
    when (state.status()) {
        AssetPackStatus.DOWNLOADING -> {
            val percent = 100.0 * state.bytesDownloaded() / state.totalBytesToDownload()
        }
        AssetPackStatus.COMPLETED -> { /* asset pack is ready to use */ }
        AssetPackStatus.FAILED -> { /* state.errorCode() */ }
    }
}
```

via `com.google.android.play:asset-delivery-ktx:2.3.0`. This is Google's own
delivery infrastructure rather than a model file sitting behind a plain
download URL, which is the main reason to reach for it over a first-run
download once a model is large enough that install-time bundling in the base
module no longer fits: it gets progress reporting, resume, and Play-managed
storage for close to free, where a first-run download has to build all of
that by hand.

None of this changes the underlying rule: most fine-tuned chat models at a
usable quantisation still land past the point where bundling of any kind
makes sense, on either platform. The Android-specific point is narrower:
when bundling genuinely is the right call, for example a small model and a
hard requirement to work fully offline from first launch, Android's ceiling
for that call is far higher than iOS's, because of asset packs specifically.

## Platform gotchas

- **Treat a CPU fallback path as mandatory, not optional.** Android device
  fragmentation is real in a way iOS's narrower hardware matrix is not: chip
  vendor, GPU driver quality, and available acceleration backends vary
  enormously across the installed base. llama.cpp's own official Android
  example builds CPU-only by default (ARM's KleidiAI kernels plus OpenMP
  threading, confirmed by reading that example's own CMake configuration),
  with no GPU backend enabled out of the box. Ship a working CPU path first
  and treat any GPU backend as a targeted optimisation layered on top for
  the subset of devices confirmed to run it well, not the baseline the app
  is built around.
- **NNAPI is deprecated; do not build a new integration around it.** Google's
  own NDK migration guide states it plainly: "The Neural Networks API (NNAPI)
  is deprecated. It was introduced in Android 8.1 to provide a unified
  interface for hardware accelerated inference for on-device machine
  learning, and deprecated in Android 15." Google's stated replacement path
  is TensorFlow Lite in Play Services or LiteRT's own delegates (GPU, and
  emerging NPU delegates) for hardware acceleration, and AICore for GenAI
  foundation models specifically. An existing NNAPI integration still runs,
  but a new one should target a LiteRT delegate instead.
- **GPU backend selection differs by stack.** MediaPipe's LLM Inference API
  exposes a CPU-or-GPU backend preference on Android specifically, per
  Google's own guide, though the exact builder method was not confirmed
  against a primary source in this pass; check the current
  `LlmInferenceOptions.Builder` reference before relying on the exact call.
  llama.cpp's own CMake build supports `-DGGML_VULKAN=ON` for a broadly
  portable GPU backend and `-DGGML_OPENCL=ON` specifically for GPU
  acceleration on recent Adreno hardware, per llama.cpp's own build
  documentation; neither is on by default in the official Android example,
  so enabling either is a build-configuration change to make deliberately,
  not something that happens automatically.
- **MediaPipe LoRA support is narrow.** Per Google's own LLM Inference
  guide, LoRA works "for all Gemma variants and Phi-2 models for the GPU
  backend, with LoRA weights applicable to attention layers only." A LoRA
  adapter for another model family, or a request to apply one on the CPU
  backend, is not supported.
- **The llama.cpp NDK path does not expose runtime LoRA in its official
  Kotlin wrapper, even though the underlying C API supports it.**
  llama.cpp's `llama-server` and llama.rn both expose runtime LoRA loading;
  the official Android example's Kotlin interface, read directly from its
  source, offers `loadModel`, `setSystemPrompt`, `sendUserPrompt`, `bench`,
  `cleanUp`, and `destroy`, with no LoRA method. Adding one means extending
  the JNI layer, not calling something already there.
- **64-bit versus 32-bit is a real split, not a formality.** The official
  llama.cpp Android example only builds for `arm64-v8a` and `x86_64`, per
  its own Gradle configuration, leaving older 32-bit ARM devices
  unsupported by that specific example. Check whichever wrapper is chosen
  for its own ABI coverage before assuming a build works everywhere the app
  otherwise supports.
- **Quantisation floor for GGUF on mobile.** Below roughly 3B parameters,
  quantising past Q4_K_M is a known quality cliff, not a bug to debug later.
  Start at Q4_K_M and only go lower after confirming quality still holds.
- **Neither MediaPipe, LiteRT-LM, nor a llama.cpp-based wrapper solves
  per-device model versioning.** If the fine-tune changes, a bundled asset,
  a Play asset pack, and a server-hosted download URL all need an explicit
  update path; none of these runtimes pushes a new model version to an
  installed app on their own.

Full depth on each path, including the `litert-torch` conversion flags,
`.task` versus `.litertlm` and the pre-converted model set, and the GGUF
conversion and quantisation commands, is in
`references/mediapipe-litert-path.md` and `references/llamacpp-jni-path.md`.

## Hand off to

- The artifact shape is not yet confirmed, or it is a merged checkpoint or a
  PEFT adapter that needs converting first: **inspecting-a-model-bundle**
- The model loads but generates badly, repeats, never stops, or the adapter
  will not apply: **debugging-a-bad-fine-tune**
- It is not yet known whether the fine-tune is actually better than the base
  model it started from: **evaluating-a-tuned-model**
- The question is whether running this on device is actually cheaper than an
  API at the expected usage level: **costing-a-model-vs-an-api**
- The target platform is React Native or Expo, native iOS, or Flutter
  instead of native Android: **shipping-a-model-in-a-react-native-app**,
  **shipping-a-model-in-an-ios-app**, or
  **shipping-a-model-in-a-flutter-app**

Files in this skill

  • SKILL.md17.7 KB
  • references/llamacpp-jni-path.md6.5 KB
  • references/mediapipe-litert-path.md5.5 KB
  • references/model-delivery-and-size-budgets.md4.9 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…