Put a local language model inside a React Native or Expo app and get it generating on device. Covers react-native-executorch for .pte models and llama.rn for GGUF models, choosing between them, native build configuration and model loading, streaming answers into the UI as tokens arrive, and whether to bundle the model in the binary or download it on first run. Use when the project is React Native or Expo and someone wants on-device or offline AI, local inference, or a model running without an...
Installs into .claude/skills of the current project.
Are you the author of Shipping A Model In A React Native App?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/ertasai-shipping-a-model-in-a-react-native-app)
---
name: shipping-a-model-in-a-react-native-app
description: >-
Put a local language model inside a React Native or Expo app and get it
generating on device. Covers react-native-executorch for .pte models and
llama.rn for GGUF models, choosing between them, native build configuration
and model loading, streaming answers into the UI as tokens arrive, and
whether to bundle the model in the binary or download it on first run. Use
when the project is React Native or Expo and someone wants on-device or
offline AI, local inference, or a model running without an API. Not for iOS
native Swift projects, not for Android native Kotlin projects, not for
Flutter, and not for models running on a server.
license: Apache-2.0
metadata:
version: "0.1.0"
author: "Edward Xi Yang, Ertas AI"
stage: "ship"
previous_skill: "inspecting-a-model-bundle"
---
# Shipping a model in a React Native app
Two runtimes cover React Native and Expo, and they do not take the same
artifact. Picking between them is the first decision here, before any
package gets installed, because reading the wrong runtime's docs for an hour
is a common way to lose an afternoon.
## Which artifact shape this needs
| You are holding | Package | What it eats |
|---|---|---|
| A `.pte` file, plus `tokenizer.json`, `tokenizer_config.json`, `config.json` | `react-native-executorch` | An ExecuTorch program |
| A single `.gguf` file | `llama.rn` | GGUF, self-contained |
These are two separate native modules with two separate build steps. There is
no package that takes both. If a `.pte` file gets handed to `llama.rn`, or a
`.gguf` file to `react-native-executorch`, it fails to load, full stop.
**If what you are holding is neither of these,** a merged Hugging Face
checkpoint (`config.json` + `model*.safetensors`) or a PEFT adapter directory
(`adapter_config.json` + `adapter_model.safetensors`), it is not shippable
into React Native as is. Run **inspecting-a-model-bundle** first to confirm
which shape you actually have, then convert:
- To `.pte`: `optimum-cli export executorch`, covered below and in
`references/executorch-path.md`.
- To GGUF: `convert_hf_to_gguf.py` then `llama-quantize`, covered below and
in `references/llama-rn-path.md`.
A PEFT adapter directory converts to neither format directly. Either merge it
into its base model first and export from the merged checkpoint, or, for the
GGUF path only, convert the adapter itself with `convert_lora_to_gguf.py` and
load it as a runtime LoRA against a base GGUF (see the streaming and LoRA
notes below).
**If it is genuinely unclear which of the two to pick** and both are
available, GGUF is the safer default: `llama.rn` is the only mainstream
mobile runtime that takes a plain GGUF with no conversion step, it supports
hot-swappable runtime LoRA adapters, and it has no bundle-size ceiling of its
own beyond the app store limits below. `react-native-executorch` is the
better pick when CoreML acceleration on Apple GPUs is wanted specifically, or
when the model was already exported to `.pte` for another ExecuTorch target.
### If ExecuTorch is the answer, install Software Mansion's skill as well
Software Mansion write `react-native-executorch`, and they ship an agent skill
for it in their own repo:
```bash
npx skills add software-mansion/react-native-executorch
```
Theirs targets the current published package API and covers every public hook,
including the vision, speech, embedding and tokenizer hooks this skill leaves
alone. Where the two overlap on `useLLM`, treat theirs as authoritative,
because it ships with the package and moves when the package moves.
The division of labour worth keeping in mind: this skill covers the choice
between `react-native-executorch` and `llama.rn`, the artifact conversion each
one needs, and the delivery and size budget that follows from it. Theirs covers
what to call once `react-native-executorch` is the runtime.
## Install and wiring
### react-native-executorch (.pte)
```bash
npm install react-native-executorch
# Expo managed workflow also needs the resource fetcher:
npm install react-native-executorch-expo-resource-fetcher
```
Version 0.9.2 (2026-06-17), nightly channel at 0.10.0. Minimum iOS 17.0,
minimum Android 13. **New Architecture only.**
Wiring, from the package README:
```tsx
import {
useLLM,
models,
initExecutorch,
} from 'react-native-executorch';
import { ExpoResourceFetcher } from 'react-native-executorch-expo-resource-fetcher';
initExecutorch({
resourceFetcher: ExpoResourceFetcher,
});
```
`initExecutorch` runs once, before any component calls `useLLM`. The
`resourceFetcher` is only needed under Expo; a bare React Native project can
skip it.
### llama.rn (GGUF)
```bash
npm install llama.rn
```
Version 0.12.7 (2026-07-22). **New Architecture required** since v0.10.
For Expo, add the config plugin to `app.json`:
```json
{
"expo": {
"plugins": [
[
"llama.rn",
{
"enableEntitlements": true,
"forceCxx20": true,
"enableOpenCL": true
}
]
]
}
}
```
`enableOpenCL` is only relevant on Android; leave it on unless targeting iOS
exclusively.
Both packages require the New Architecture flag turned on for the project
(`newArchEnabled=true` in `gradle.properties` on Android, the equivalent Expo
or bare RN setting on iOS) before either will build. This is not optional for
either runtime.
## A minimal working example
### react-native-executorch: load, generate, stream
```tsx
import {
useLLM,
models,
Message,
initExecutorch,
} from 'react-native-executorch';
import { ExpoResourceFetcher } from 'react-native-executorch-expo-resource-fetcher';
initExecutorch({ resourceFetcher: ExpoResourceFetcher });
function Chat() {
const llm = useLLM({ model: models.llm.lfm2_5_1_2b_instruct() });
const send = async () => {
const chat: Message[] = [
{ role: 'system', content: 'You are a helpful assistant.' },
{ role: 'user', content: 'What is the meaning of life?' },
];
await llm.generate(chat);
// llm.response holds the full text once generate() resolves
};
// llm.response updates as tokens arrive, so render it directly for
// streaming: it re-renders the component on every new token, this is not
// a separate callback API the way llama.rn's is.
return null; // wire llm.response, llm.isGenerating, llm.isReady into UI
}
```
`llm.isReady` reports whether the model finished loading, and
`llm.isGenerating` reports whether a `generate()` call is in flight. Check
`isReady` before calling `generate()`.
### llama.rn: load, generate, stream
```js
import { initLlama } from 'llama.rn';
const context = await initLlama({
model: 'file:///path/to/model.gguf',
use_mlock: true,
n_ctx: 2048,
n_gpu_layers: 99, // Metal / OpenCL offload; 99 offloads everything that fits
});
const result = await context.completion(
{
messages: [{ role: 'user', content: 'What is the meaning of life?' }],
n_predict: 256,
},
(data) => {
// called once per token as it streams in
console.log(data.token);
},
);
console.log(result.text); // full text, once generation finishes
console.log(result.timings);
```
`n_ctx` is the context window in tokens; `n_gpu_layers` controls how many
transformer layers offload to Metal on iOS or OpenCL on Android, `99` is
shorthand for "offload as many as fit." Streaming here is an explicit
per-token callback, not a state variable that re-renders on its own, the
opposite of the `useLLM` pattern above.
**Runtime LoRA, llama.rn only:** a GGUF LoRA adapter, produced by
`convert_lora_to_gguf.py`, can be applied and swapped at runtime:
```ts
await context.applyLoraAdapters([{ path: '/path/to/adapter.gguf', scaled: 1.0 }]);
await context.removeLoraAdapters();
```
`react-native-executorch` has no equivalent runtime adapter API. Any LoRA
used with it has to be merged into the base model before export, which means
a new `.pte` export for every adapter variant, not a hot swap.
## The delivery decision
Once the model runs, the next question is how it reaches the device: bundled
into the app binary, or downloaded on first run. Full detail, the app store
size ceilings, and the first-run download experience are in
`references/model-delivery-and-size-budgets.md`.
The one React Native specific number that overrides the general table:
`react-native-executorch`'s `require()` route for bundling a `.pte` from the
assets folder has a **hard 512 MB cap**, documented by the package itself,
independent of what the app stores would otherwise allow. Past that size,
`react-native-executorch` has two other supply routes: a remote URL fetched
at runtime, or a `file://` path when the user supplies their own model file.
`llama.rn` has no equivalent package-level cap; its ceiling is whatever the
app stores and install-conversion tolerance allow.
For most fine-tuned chat models at a usable quantisation, the artifact lands
well past the point where bundling makes sense on either package. Download on
first run is the default worth reaching for; treat bundling as the exception
that needs a specific justification, not the starting assumption.
## Platform gotchas
- **New Architecture is mandatory for both packages**, not a performance
option. A project still on the legacy architecture will not build against
either.
- **react-native-executorch backends are XNNPACK (CPU) and CoreML (Apple
GPU) only**, via `optimum-executorch`. There is no CUDA or generic GPU
backend for this path; Android acceleration beyond CPU is not available
through ExecuTorch today.
- **llama.rn Android OpenCL needs a manifest entry.** For Adreno 700-class
GPUs, add `<uses-native-library android:name="libOpenCL.so"
android:required="false" />` to `AndroidManifest.xml`, or `n_gpu_layers`
silently falls back to CPU.
- **llama.rn's Hexagon NPU path is experimental**, Snapdragon SM8450 and
later only, and needs `libcdsprpc.so` declared plus `devices: ['HTP0']`
passed at init. Treat it as a bonus, not a baseline target.
- **The `.pte` bundle is not one file.** `react-native-executorch` needs the
`.pte` alongside `tokenizer.json`, `tokenizer_config.json`, and ideally the
model's variant `config.json`. Shipping only the `.pte` and expecting it to
load is a common mistake.
- **ExecuTorch's program/weight split (`.ptd`) exists but is not the
default path.** A model can be exported as a `.pte` program plus a
separate `.ptd` weights file, which is the mechanism that in principle
lets several LoRA-adapted `.pte` files share one set of foundation
weights. Treat this as an advanced option to check against current
ExecuTorch release notes before relying on it, not a settled feature to
build a shipping plan around.
- **Quantisation floor for GGUF on mobile.** Below roughly 3B parameters,
quantising past Q4_K_M is a known quality cliff, not a bug to debug later.
Start at Q4_K_M and only go lower after confirming quality still holds.
- **Neither package solves per-device model versioning.** If the fine-tune
changes, both the bundled asset and any server-hosted download URL need an
explicit update path; there is no OTA model update built into either
runtime.
Full depth on each path, including the `optimum-executorch` export flags and
llama.rn's context configuration, is in `references/executorch-path.md` and
`references/llama-rn-path.md`.
## Hand off to
- The artifact shape is not yet confirmed, or it is a merged checkpoint or a
PEFT adapter that needs converting first: **inspecting-a-model-bundle**
- The model loads but generates badly, repeats, never stops, or the adapter
will not apply: **debugging-a-bad-fine-tune**
- It is not yet known whether the fine-tune is actually better than the
base model it started from: **evaluating-a-tuned-model**
- The question is whether running this on device is actually cheaper than
an API at the expected usage level: **costing-a-model-vs-an-api**
- The target platform is a native iOS, native Android, or Flutter app
instead of React Native or Expo: **shipping-a-model-in-an-ios-app**,
**shipping-a-model-in-an-android-app**, or
**shipping-a-model-in-a-flutter-app**
- The runtime is settled as `react-native-executorch` and the open question is
which hook to call or how to configure it: Software Mansion's own skill,
`npx skills add software-mansion/react-native-executorch`