Reuse what LM Studio already has on this computer through its `lms` CLI: adopt its running server into the person's fleet, serve its MLX models, or serve its GGUF files without it. Load before touching LM Studio, its models or its server.
1,124 stars
0 votes
0 copies
0 views
Added September 30, 2026
ai-agentsgobashrailsgitapi
Works with
cli
api
Security analysis
D42/100
criticalPipes output to a shell interpreter
mediumUses curl or wget to download content
criticalDownloads and executes remote scripts — classic supply chain attack
Installs into .claude/skills of the current project.
Are you the author of Engine Lm Studio?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/autonomous-ai-engine-lm-studio)
---
name: engine-lm-studio
description: "Reuse what LM Studio already has on this computer through its `lms` CLI: adopt its running server into the person's fleet, serve its MLX models, or serve its GGUF files without it. Load before touching LM Studio, its models or its server."
---
# LM Studio (`lms` CLI, no clicks)
Official docs, read 2026-09-29 (source files in github.com/lmstudio-ai/docs):
[App, llmster and lms](https://lmstudio.ai/docs/app/basics/lmstudio-vs-llmster-vs-lms) ·
[Headless](https://lmstudio.ai/docs/developer/core/headless) · [lms server start](https://lmstudio.ai/docs/cli/serve/server-start) ·
[lms server stop](https://lmstudio.ai/docs/cli/serve/server-stop) · [lms load](https://lmstudio.ai/docs/cli/local-models/load) ·
[lms import](https://lmstudio.ai/docs/cli/local-models/import) · [lms daemon up/down](https://lmstudio.ai/docs/cli/daemon/daemon-up) ·
[Idle TTL](https://lmstudio.ai/docs/developer/core/ttl-and-auto-evict) · [OpenAI compatibility](https://lmstudio.ai/docs/developer/openai-compat) ·
[Tools](https://lmstudio.ai/docs/developer/openai-compat/tools) · [Structured output](https://lmstudio.ai/docs/developer/openai-compat/structured-output) ·
[Import models](https://lmstudio.ai/docs/app/advanced/import-model) · [Server settings](https://lmstudio.ai/docs/developer/core/server/settings)
Tested: LM Studio 0.4.25 (`brew install --cask lm-studio`), MacBook Pro M1 Pro 32 GB, macOS 26.6, 2026-09-29.
Tags: `[doc]` official page above, `[run]` seen on the tested machine, `[?]` unverified.
## When to use it
- **A model in LM Studio's folder**, GGUF or MLX (`fleet models` → START WITH `lm-studio`): LM Studio runs
it, because it downloaded it and loads it.
- **Its server answering**: adopt it, but load the model yourself with a 64K+ context (below) — a model
it loads on demand gets 8192 [run], too small for an agent.
- **`(start it)`**: start its server yourself (below); an LM Studio that is off is the normal case, not
a reason to switch engines.
- **An MLX model elsewhere and mlx-lm not installed**: LM Studio runs MLX on Apple silicon (its
structured output uses Outlines for MLX [doc]), so use it instead of installing anything.
- Never install LM Studio just to serve a model: Grid's engine covers GGUF, mlx-lm covers MLX.
## Where its models live / how to list them
- `~/.lmstudio/models/<publisher>/<model>/<file>` — it keeps Hugging Face's layout [doc]. The folder can be
moved in the app's My Models tab [doc]; no documented setting file or env var [?]. Older versions used
`~/.cache/lm-studio/models` [?]. `fleet models` scans both.
- `lms ls --json` → `modelKey`, `path`, `format` (gguf/mlx), `sizeBytes` [run]. Keys prefix-match: a short
key matched two downloaded models and loaded "the first one" [run] — always pass the exact key.
- A fresh install holds only an embedding model, no chat model [run].
## Installed? Running?
- App: `/Applications/LM Studio.app`. `lms` lives in `~/.lmstudio/bin` and appears only after the app ran
once [run]. Headless servers use **llmster** instead of the app: `curl -fsSL https://lmstudio.ai/install.sh | bash` [doc].
- `lms server status` → "The server is not running." or its port [run]. Launching the app does not
start the server [run]. `fleet models` lists a running one as `kind: lm-studio` (its `owned_by` is
`organization_owner` [run]).
## Start an already-downloaded model
Port P from outside `machine.listeningPorts`. Always `--port`: without it the last used port is reused [doc].
export PATH="$HOME/.lmstudio/bin:$PATH"
lms server start --port P --bind 127.0.0.1 # starts LM Studio in the background if needed [doc][run]
lms load <exact key> --estimate-only --context-length 65536 --gpu max -y # fits? [doc][run]
lms load <exact key> --context-length 65536 --gpu max --identifier <id> -y
- With the app closed, `server start` printed "Waking up LM Studio service..." and was up in 4 s; it runs
the app as `LM Studio --run-as-service`, with no window [run]. It returns at once (not a foreground process).
- `--estimate-only` prints GPU and total memory for that context and whether guardrails allow it [doc][run].
- A model loaded with `lms load` has no idle timer and stays until unloaded [doc]; on-demand loads unload
after 60 minutes idle [doc]. Nothing downloads on load [run].
## Ready means
Run `"$GRID_FLEET" verify --at http://127.0.0.1:P/v1 --model <id> --kind lm-studio` — it performs the
request checks with deadlines and prints each one. Also:
1. `lms server status` shows port P [run]. 2. `GET http://127.0.0.1:P/v1/models` lists `<id>` [doc][run].
3. LM Studio reports `<id>` loaded with ≥ 65536 (`/api/v0/models`, the CONTEXT `lms ps` shows) [run]. 4. One bounded `/v1/chat/completions` request with
`model: <id>`, `max_tokens` 16 → non-empty `content` [run].
A 401 means the person turned on "Require Authentication" [doc]: do not change their setting; say so and
serve the file with Grid's engine instead.
## Join Harness Compute
"$GRID_FLEET" run -- join GRID --at http://127.0.0.1:P/v1 -m <id> --advertise-as ALIAS
`/v1` is required: with the bare root, Grid's capability probe records structured output as unsupported [run].
Grid's own detector finds LM Studio only on 1234 [run].
## Tool calls, JSON output, thinking
- Tools: `tools` on `/v1/chat/completions` and `/v1/responses`, OpenAI format [doc]. Grid's probe: tools yes [run].
- Structured output: `response_format` with `json_schema` [doc]; GGUF uses llama.cpp grammars, MLX uses
Outlines [doc]. Grid's probe: json_schema yes, json_object no [run].
- Thinking control through the OpenAI endpoint: not documented [?]. Check `content` in the ready request.
## Memory and speed knobs
`--context-length` (≥ 65536 here), `--gpu max|off|0–1` (automatic when omitted) [doc], `--ttl <seconds>` [doc],
`--parallel <n>` (default 4 on this build [run]), `--speculative-draft-mtp` for models with MTP layers [run: in help].
## Stop
- Your model: `lms unload <id>` [doc]. Your server: `lms server stop` [doc] — the background service keeps running [run].
- The service you woke: `lms daemon down` stops llmster only, not the app [doc]; the app service ends with
`osascript -e 'quit app "LM Studio"'` [run].
- If LM Studio was running before you came, unload only your model and leave the rest as it was.
## Known failures → what to do
| Sign | Do |
|---|---|
| "2 models match the provided model key" | use the exact `modelKey` from `lms ls --json` [run] |
| CONTEXT 8192 in `lms ps` | the model was loaded on demand; unload it and `lms load` with `--context-length 65536` [run] |
| `lms get … -y` shows a spinner forever | it downloads in the background [run]; never start a download without the go-ahead |
| a foreign GGUF must appear in LM Studio | `lms import -l -y --user-repo <user>/<repo> FILE.gguf`: **without `-l` it moves the file** [doc]; the file name must end in `.gguf` and the link must point at the real file, not a temp copy [run] |
| 401 from `/v1/models` | authentication is on in the person's settings [doc]; do not turn it off |
| a model already on disk is not in `lms ls` | no download: an **MLX folder** is symlinked as `~/.lmstudio/models/<publisher>/<name>` → the folder, then `lms ls` lists it [run]; a **GGUF** goes in with `lms import -l` (row above). Neither copies |
| answers come back as reasoning only; no request switch turns thinking off | some models keep thinking under LM Studio whatever the request says [run]. That is not a failure: `fleet verify` asks once more with room to think and passes, noting it. Never hand-test switches with `curl` (it cost minutes [run]); tell the person replies will be slower, and offer a model that can switch it off if they want that |