Pitfalls for writing/editing marimo notebook .py files, and for mutating a RUNNING notebook via marimo-pair code_mode (ctx.edit_cell/create_cell/move_cell). Covers: import patterns (never alias marimo as mo), cell params (mo received as a parameter), App() init, unique-variable rules incl. the loop-variable cell-collision (namespace loop vars per cell), cross-cell symbol removal, ran-error traceback triage, import-cell return values, underscore-prefix privacy, button/toast/callout/slider/prog...
Installs into .claude/skills of the current project.
Are you the author of Marimo Notebook Patterns?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/ericmjl-marimo-notebook-patterns)
---
name: marimo-notebook-patterns
description: >-
Pitfalls for writing/editing marimo notebook .py files, and for mutating a RUNNING notebook via marimo-pair code_mode (ctx.edit_cell/create_cell/move_cell). Covers: import patterns (never alias marimo as mo), cell params (mo received as a parameter), App() init, unique-variable rules incl. the loop-variable cell-collision (namespace loop vars per cell), cross-cell symbol removal, ran-error traceback triage, import-cell return values, underscore-prefix privacy, button/toast/callout/slider/progress_bar API signatures, StaleCellError, ctx.cells staleness after mutation, code-mode edit flush-on-exit, move_cell for positioning, __main__ clobber on save, no asyncio.run (top-level await), importlib.reload after library edits, hide_code rubric, plotly-over-matplotlib conversion, live training curves via anywidget+_esm, and LanceDB docstore table-name collisions. Load when authoring or editing ANY marimo notebook .py file, or when a create_cell/edit_cell batch rolls back with a multiple-definition error.
created_by: autolearn
created_at: "2026-06-09"
---
# Marimo Notebook Patterns
Common pitfalls when writing and editing marimo notebook .py files. Covers: import patterns (never alias marimo as mo), cell parameter requirements (mo must be received as parameter), and app initialization (use marimo.App() not mo.App()).
## Instructions
TODO: Add specific instructions based on observed patterns.
## import-alias-conflict
- Never use `import marimo as mo` at module level in marimo notebook files. Cell functions receive `mo` as a parameter (e.g. `def _(mo):`), and a module-level alias shadows it. Always use `import marimo` (no alias) at module level and `app = marimo.App()`.
- The import alias conflict causes marimo's static parser to fail with the error: 'Static loading of notebook failed. Please report this issue to the marimo team...' — a cryptic message that doesn't mention the import. When a notebook throws this static loading error, check for `import marimo as mo` at module level as the first suspect. Also check for duplicate marimo imports and missing `__generated_with` header lines, which are common co-occurring issues when notebook files are manually created or edited.
## cell-parameter-verification
- After editing or creating marimo notebooks, verify that every cell function using `mo.md()`, `mo.ui.*`, or any `mo.*` call actually receives `mo` as a function parameter. A cell like `def _(): return mo.md('text')` will fail at runtime because `mo` is not in scope. Fix to `def _(mo): return mo.md('text')`.
## debugging-loop
- When fixing a marimo notebook that should 'run cleanly from top to bottom', execute all cells from top to bottom to observe actual runtime errors, fix them, and re-run. Don't statically analyze file structure, check headers, or hypothesize about kernel state — runtime errors reveal what static analysis misses. The correct loop is: run all cells → observe errors → fix → re-run. Use marimo-pair or the running kernel to execute cells.
## unique-variable-rule
- In marimo, each variable name must be defined in exactly one cell across the entire notebook. You cannot have or any other variable definition in multiple cells — marimo treats each variable as a reactive reference that must have a single source cell. To share dependencies across cells, pass them as function parameters (e.g. ). After editing notebooks, always run to catch multiply-defined variables.
- In marimo, each variable name must be defined in exactly one cell across the entire notebook. You cannot have df_plot or any other variable defined in multiple cells — marimo treats each variable as a reactive reference that must have a single source cell. To share dependencies across cells, pass them as function parameters (e.g., def _(df_source):). After editing notebooks, always run all cells to catch multiply-defined variables. COMMON TRAP: When you want to rename a column or transform a dataframe, do NOT create a new cell that redefines the variable (e.g., df_vari = df_vari.rename(...)). This creates a multiply-defined error because the original cell already defines df_vari. Instead, modify the original defining cell to include the rename, or use a different variable name for the result.
- DUPLICATE-DEFINITION DIAGNOSTIC GOTCHA (the expensive one): when a name (Counter, tool, payload, or any import/assignment target) is defined in TWO cells, marimo raises MultipleDefinitionError on the SECOND defining cell and marks it 'marimo-error' (output suppressed). The symptom downstream is a NameError on that name in cells that reference it — and because the cell WIRING looks correct (consumer signatures reference the name properly), the NameError READS like a broken reactive graph / a stale dependency edge, NOT like a duplicate-definition error. This misdirection is the single biggest time-sink: agents recreate cells, rewire the graph, restart the server, and chase a 'graph bug' that is actually a MultipleDefinitionError one cell over. DIAGNOSTIC PROCEDURE (do this FIRST): when a cell's status is 'marimo-error', read cell.output.data for the cause — e.g. MultipleDefinitionError(name='Counter', cells=('Kclp',)) tells you the offending name AND the other cell ID that defines it — BEFORE assuming a graph-wiring bug and recreating cells. FIX: (1) centralize shared IMPORTS in the 'with app.setup:' block (Counter, tool, json, dedent, etc.) so they are globally visible to all cells without per-cell imports — never 'from X import Y' the same Y in two cells; (2) never reuse a LOCAL variable name (e.g. payload) across two cells — merge the logic or rename. Distinct from reactive-graph-registration-corruption (that cell runs WITHOUT error yet its output is unregistered; MultipleDefinitionError shows an explicit marimo-error status). Sibling of code-mode-edit-flush-on-exit (stale READ-BACK within one block) and ctx.cells-stale-after-mutation (stale ctx.cells after delete/reorder) — all three masquerade as graph bugs but have distinct root causes.
- MATPLOTLIB PLOTTING-CELL COLLISION (the most common multi-cell unique-var violation, alphaxiv-marimo-competition-submission 07-07): every matplotlib plotting cell instinctively opens with 'fig, ax = plt.subplots()' — but BOTH 'fig' and 'ax' are assignment targets, so the SECOND plotting cell that does the same hits MultipleDefinitionError ('fig is defined in multiple cells') and marimo rolls back the ENTIRE batch (the new cells are NOT created — '__aexit__' error during validation). This is the plotting-cell analog of the 'mo' alias / Counter / json import collisions already documented, but EASIER to hit because the 'fig, ax = plt.subplots()' idiom is muscle memory and a single notebook commonly has 3-5 plotting cells (frontier plot, heatmap, race curves, etc.). FIX: give every plotting cell a DISTINCT per-cell figure/axes NAME — namespace by purpose, e.g. 'frontier_fig, frontier_ax', 'heat_fig, ax0, ax1, ax2', 'race_fig, race_ax'. The last expression must then be the namespaced fig (e.g. 'return heat_fig') so marimo renders it. Do NOT use the bare 'fig'/'ax' names in any plotting cell if another cell already defines them, and do NOT reuse throwaway '_fig'/'_ax' (those collide too if more than one cell uses them). Going forward when adding a plotting cell to a notebook that ALREADY has plots, scan the existing cells for 'fig =' / 'ax =' first and pick an unused prefix. Idiomatic alternative that sidesteps naming entirely: assign the figure to a unique name AND set it as the cell's return, OR restructure so each plot lives in its own clearly-named cell function. Distinct from import-alias collisions (those are MODULE imports; this is matplotlib RESULT object names) and from the df_vari rename trap (that's redefining a DATA variable; this is the standard plt result objects).
- CONTEXT-MANAGER TARGET COLLISION (build-deep-research-agent notebook 01, 07-10): a 'with urlopen(request) as response:' in one cell and 'response = bot(...)' in another cell collide because the context-manager binding target ('response' in 'with ... as response:') is a cell-level ASSIGNMENT TARGET in marimo's reactive system — it is NOT scoped to the with-block. This is the same class as the loop-variable collision (i/m/b) and the matplotlib fig/ax collision, but easier to miss because (a) the name looks block-local, (b) common names like 'response', 'f', 'file' are natural to reuse across cells, and (c) 'with X as Y' is muscle memory that doesn't trigger 'this is a reactive name' alarm. FIX: namespace the context-manager target per cell with a DESCRIPTIVE unique name (e.g. 'ping_resp', 'api_response', 'config_file') to avoid collision with any other cell's assignment. Do NOT use underscore-prefix names (_resp, _file_handle) — the user prefers no underscore-prefix/'private' variables (autolearn 07-10: 'no private variables either too, please'). When diagnosing a MultipleDefinitionError, grep for 'with .* as ' patterns alongside '=' assignments and 'for ' targets — all three are assignment targets. Discovered build-deep-research-agent notebook 01 startup_validation cell colliding with ex1_run cell on the name 'response'.
## expression-in-try-block
- In marimo notebooks, bare expressions inside try/except blocks are still captured by marimo for display. For example, `G.edges[15, 16]` in a try block will cause marimo to attempt displaying it. Assign to a variable: `_ = G.edges[15, 16]` to suppress display, even inside error-handling code where you expect the expression to raise an exception.
- RESOLVED CONTRADICTION (07-10): an earlier version of this section claimed bare expressions inside try/except blocks ARE captured by marimo for display. That is UNRELIABLE — two independent observations on 07-10 (memory #604 startup_form cell + Part 4 ex1_run cell) show that mo.md()/mo.vstack() expressions inside try-except blocks are DISCARDED (not rendered), exactly like if/for/with blocks. The ONLY reliable rule: marimo renders the cell's LAST TOP-LEVEL expression (or explicit return value); expressions inside ANY nested block (if/for/with/try/except) are STATEMENTS that evaluate and are discarded. Treat the old 'try/except ARE captured' claim as superseded. BARE-RETURN COMPOUNDING FACTOR (Part 4 ex1_run, 07-10): a cell that uses mo.md() as a 'print-like' statement inside try/except for status+result/error display, ending with a bare 'return' (returning None), renders NOTHING — the return value (None) is the cell output, overriding any earlier expressions. COMMON PATTERN that triggers this: 'mo.md("Running...")' then 'try: result = agent(...); mo.md(str(result)) except Exception as e: mo.md(f"Error: {e}")' then bare 'return'. FIX: capture the desired output into a single variable and make it the cell's last expression: 'output = mo.md("Running..."); try: ...; output = mo.md(str(result)) except: ...; output = mo.md(f"Error: {e}"); output' (or 'return output'). Do NOT use mo.md() as a side-effect print statement — it returns an Html object that only displays as the cell's final expression.
## import-cell-return
- In marimo, a cell that imports a library must explicitly return the imported symbol for it to be available to other cells. For example, an import cell doing `import matplotlib.pyplot as plt` must return `(plt,)` at the end of the cell function — marimo's reactive system only propagates variables that are returned by their defining cell. If a downstream cell gets 'X is not defined' despite an import cell existing, check that the import cell actually returns the imported name. This is distinct from the cell-parameter pattern (where mo must be a parameter) — this is about the cell's return value enabling cross-cell variable propagation.
## code-mode-variable-injection
- When editing cells through marimo-pair's code_mode (ctx.edit_cell, ctx.create_cell), the cell function parameter mechanism from the .py file does NOT apply. Cells execute as raw code — there is no def _(mo): wrapper. Variables like mo must either: (1) be imported directly in the cell (import marimo as mo), or (2) be available via the reactive graph (defined and returned by another cell). This is a common source of confusion when switching between file editing and code_mode editing. The cell-parameter-verification rule (mo must be a parameter) only applies to the .py file on disk, not to code_mode cell edits.
- MULTIPLY-DEFINED trap (the 'mo' alias): `import marimo as mo` COUNTS AS DEFINING the name `mo` (it is an assignment target), so marimo's unique-variable check fires with 'Multiply-defined names: mo' if you place it in MORE THAN ONE code_mode cell — the same constraint as any other variable (see unique-variable-rule), but easy to hit because the instinct is to import mo in every cell that uses it. The safe pattern: exactly ONE cell does `import marimo as mo` (the imports cell); other cells reference bare `mo` (it propagates via the reactive graph per option 2 above — the imports cell defines+returns it, consumers receive it as a reactive dependency) OR use `marimo.` after a non-aliased `import marimo`. Do NOT sprinkle `import marimo as mo` into every code_mode cell 'to be safe' — that is the multiply-defined trap. Contrast with .py-file notebooks where `mo` is auto-injected as a cell parameter (`def _(mo):`) so no import is needed in any cell; code_mode has no such injection — hence the single imports cell. Discovered alphaxiv-marimo-competition-submission (07-07): batch-creating several cells each starting with `import marimo as mo` hit 'Multiply-defined names: mo'; fix was one imports cell owning the alias and consumer cells referencing it bare.
## named-cells
- Every cell in a marimo notebook must have a descriptive name. Use @app.cell(name='descriptive_name') instead of bare @app.cell or unnamed def _(). Named cells make the notebook structure navigable and self-documenting.
## markdown-sandwich
- Every code cell in a marimo notebook must be sandwiched between markdown cells. A markdown cell should precede each code cell to describe intent (what the code does and why), and a markdown cell (or the next section's markdown) should follow to provide context for results or transitions. This creates a readable, narrated notebook.
## button-api-signature
- In marimo, mo.ui.button() first positional argument is on_click (a callable), NOT label. The label parameter is keyword-only. Writing mo.ui.button('Click me', on_click=handler) silently passes the string 'Click me' as the on_click callback, which fails at runtime. Correct usage: mo.ui.button(label='Click me', on_click=handler). When using buttons with on_click handlers, always pass both label and on_click as keyword arguments to avoid this silent positional-arg swap.
## external-editing-stale-state
- When editing marimo notebook .py files externally (via agent/editor) while marimo edit --watch is running, the file watcher detects changes but doesn't fully sync cell structure with the frontend (known bug marimo#8421). The frontend keeps stale in-memory state, causing cells to appear in error even when the file on disk is correct. Workaround: close the marimo session before editing externally, then reopen. If cells show unexpected errors after external edits, reload the notebook rather than trying to fix individual cells.
## Widget consolidation
- When building interactive anywidget components in marimo, put the widget class, initialization, navigation controls, and display in a SINGLE cell. marimo.state does not work reliably across separately-created cells. Splitting stateful UI across multiple cells causes reactivity failures and variable tracking issues. Do not create separate cells for widget definition, state management, and display — combine them.
## ctx.cells-stale-after-mutation
- When using marimo-pair (code mode), after calling or other structural mutations, becomes stale. Accessing raises 'Cell not found' even though the ID appears in available IDs. The fix: collect all cell IDs and names BEFORE any deletions, then batch-delete without re-accessing ctx.cells in between. The error message reads: 'IDs are stable across reorders — re-read ctx.cells if the notebook structure changed.' If you must access cells after a mutation, re-acquire the context first.
## code_mode cell positioning API
- When mutating a running notebook via `marimo._code_mode`, cell POSITIONING is asymmetric across the three structural calls (verified vs `AsyncCodeModeContext` sigs, marimo >=0.13): (1) `create_cell(code, *, before=None, after=None, hide_code=True, ..., name=None)` accepts `before`/`after` so a NEW cell can be placed relative to an existing one. (2) `edit_cell(target, code=None, *, hide_code=None, ..., name=None)` does NOT accept `before`/`after` -- it only rewrites the body/attributes and CANNOT move the cell; its `name=` kwarg RENAMES the cell, not reposition it. (3) To reposition an EXISTING cell, call `ctx.move_cell(target, *, before=None, after=None)` separately (chain it after an edit_cell when you both rewrote and want to move). Pitfall: reaching for `edit_cell(..., after=...)` to relocate a cell silently fails / errors with an unexpected kwarg. Discovered build-deep-research-agent Part 3 notebook restructure (07-02).
## main-block-clobber-on-save
- When marimo re-saves a notebook .py file (which it does automatically on any code_mode edit_cell/create_cell/run_cell, or any browser edit), it REGENERATES the file from its cell graph and CLOBBERS any custom `if __name__ == "__main__":` block, replacing it with marimo's DEFAULT `app.run()`. This silently destroys dual-mode entry points — e.g. a notebook that is ALSO a headless server (`if __name__ == "__main__": build_zotero_research_server().run(transport="stdio")`) loses its server startup and silently becomes a plain marimo app. Symptom: `uv run notebooks/foo.py` starts the marimo server instead of the MCP/CLI/stdio server the notebook was built to also serve. The custom __main__ is NOT a marimo cell — it is a standalone block below the cells, so marimo's re-serialization discards it. CRITICAL when the notebook backs EARS specs for dual-mode runtime (e.g. EMCP-RUN-021, EMCP-SRV-050). Prevention/recovery: (1) AVOID triggering a re-save of a dual-mode notebook via code_mode/file edits unless you re-verify the __main__ block afterward; (2) after ANY edit to a dual-mode notebook, READ the disk file's __main__ block and RESTORE the custom server-startup code if marimo replaced it with app.run(); (3) consider putting the dual-mode server startup behind a non-__main__ sentinel (e.g. an env var check inside app.setup or a separate launcher module) so it is not in the file region marimo regenerates. The disk __main__ is the ONLY place the dual-mode contract lives — audit it after every save. Discovered build-deep-research-agent Part 3 notebook (03_tools_mcp_zotero.py) 07-02: code_mode edits caused marimo to overwrite `build_zotero_research_server().run(transport="stdio", show_banner=False)` with `app.run()`, breaking the stdio MCP server.
- RECOMMENDED PERMANENT FIX (user-validated 07-02, build-deep-research-agent Part 3): move the server-mode / dual-mode entry OUT of the marimo-managed notebook file entirely into a SEPARATE shim script (e.g. scripts/serve_zotero_mcp.py whose body is just build_zotero_research_server().run(transport='stdio')). The notebook becomes purely narrative (__main__ = marimo default app.run()); the shim is the server entrypoint ('uv run scripts/serve_X.py'). marimo can never regenerate a file it doesn't own, so the clobber is impossible. Preferred over the fragile alternatives (stop-session-to-lock; leave-running-and-re-apply). When adopting the shim, update the governing specs that referenced the old single-file command: EARS entrypoint-command fields (e.g. EMCP-RUN-001 path), mode-detection wording (EMCP-RUN-020 changes from '__name__ check in one file' to 'which entrypoint is run'), LLD dual-mode section, HLD, and notebook prose (intro/recap/ex4_header) that claims 'same file dual-mode'.
## code-mode-triple-quote-collision
- When constructing a cell's CODE as a Python string to pass to ctx.create_cell / ctx.edit_cell via execute-code.sh (marimo-pair code_mode), do NOT wrap it in triple double-quotes ("""...""") if the cell body itself contains triple double-quotes — markdown cells almost always do, because they use mo.md(dedent("""...""")). The inner """ terminates the outer string prematurely, and the trailing content is parsed as code, producing a misleading 'IndentationError: unexpected indent' (or SyntaxError) at the kernel. Fix: use triple SINGLE quotes ('''...''') as the OUTER wrapper when the cell body contains triple double-quotes. A single apostrophe inside '''...''' (e.g. 'llamabot's') is fine — only three consecutive single quotes would close it. Mirror rule: if the cell body contains triple single quotes instead, wrap outer in triple double-quotes. Pre-flight check before sending: scan the cell body for the same triple-quote delimiter you chose for the outer wrapper; if it appears, switch the outer delimiter. This is the Python-string-literal-nesting analog of shell-escaping gotchas and recurs whenever you build markdown/mo.md cells via code_mode. Discovered build-deep-research-agent Part 3 (07-02): constructing an ex1 markdown cell with ex1 = """mo.md(dedent("""..."""))""" failed with IndentationError until switched to '''...'''.
- RAW-STRING ESCAPE SUBTLETY: if you wrap a markdown cell body in a RAW triple-quoted string (r"""...""" or r'''...''') and escape an embedded skeleton docstring's triple-quotes as backslash-quote-quote-quote, the escapes are PRESERVED LITERALLY (raw strings do not process backslash-escapes) — participants see the escaped quotes instead of clean triple-quotes, even though the cell status is idle (parses fine, since the skeleton is just display text inside mo.md, not executed). Fix: either use a NON-RAW string (dedent("""...""")) so the escapes evaluate to clean triple-quotes, OR (preferred) use the triple-single-quote outer wrapper recommended above so no escaping is needed at all. Discovered build-deep-research-agent Part 3 ex2 header (07-07): mo.md(r'''...''') skeleton with an escaped docstring rendered escaped quotes until switched to non-raw dedent.
- ESCALATION: cell body contains BOTH triple-double AND triple-single quotes. Common for HTML/SVG-heavy cells — mo.Html(f"""...""") uses triple-double while a nested helper function returns f'''<svg .../>''' using triple-single. Neither delimiter works as the outer wrapper (one collides with each), so the 'mirror rule' (switch delimiters) has no move to make. Fix: STOP embedding the cell code as a Python string literal in the builder. Instead write the cell code to its OWN .py file (e.g. /tmp/hero_cell_code.py containing the raw cell body with NO outer wrapper string), then in the builder script READ the file and pass it to create_cell/edit_cell: code = Path('/tmp/hero_cell_code.py').read_text(); ctx.create_cell(name='hero', code=code). This completely sidesteps Python string-literal nesting because the cell code is loaded from disk, not embedded in a literal. This is the general escape hatch when the cell body is complex enough that quoting is intractable — reach for it BEFORE burning turns trying to find a delimiter combination that avoids all collisions. Discovery path that leads here: (1) try f-string wrapper → collision; (2) try raw wrapper → still collides (raw strings prevent backslash-escaping the inner triple-quotes); (3) switch delimiters per mirror rule → still collides because the OTHER delimiter is also present; (4) file-based approach. Steps 1-3 each burn a round-trip; skip to step 4 when the cell contains mo.Html/mo.md with rich SVG/HTML art. Discovered alphaxiv-marimo-competition-submission (07-08): JiT hero cell with mo.Html(f"""...""") + jit_hero_art returning f'''<svg>...''' hit collisions on both delimiters; the file-based read() approach was the clean fix.
- - EMOJI SURROGATE-PAIR ESCAPES in anywidget ESM strings: when a cell body contains emoji (button labels, headers, decorative chars) and you encode the escape as a UTF-16 surrogate pair ('\ud83d\udcd0' for 📐 U+1F4D0), Python creates actual surrogate code points that CANNOT be UTF-8 encoded — raises 'UnicodeEncodeError: surrogates not allowed in position N' when the string is .encode('utf-8')d, written to a file, or base64-encoded. This is the same family as the triple-quote nesting issue (constructing cell code as a Python string literal that must survive transport), but on the ENCODING axis rather than the DELIMITER axis. FIX: use the single-codepoint escape with capital U + 8 hex digits ('\U0001F4D0', NOT '\ud83d\udcd0'), OR use the literal emoji character directly in the source, OR (simplest) use plain-text button labels to avoid the issue entirely. RECURRING: 2026-05-16 (anywidget voice buttons — replaced emoji with '>> Speak' / '|| Stop' plain text) and 2026-07-08 (anywidget game — removed 📐 from 'Measure target dimension' button). When a cell body with emoji fails to push via code_mode with a UnicodeEncodeError mentioning 'surrogates', the cause is always a \udXXX\udXXX surrogate-pair escape — replace with \UXXXXXXXX or plain text. Also applies to the macOS base64 wrapping gotcha: 'base64 -i' on macOS wraps at 76 chars by default, breaking base64 embedded in heredocs — pipe through 'tr -d newline' or use 'base64 -A' (no-wrap).
## code-mode-edit-flush-on-exit
- When mutating a notebook via marimo._code_mode, ctx.edit_cell / ctx.create_cell / ctx.move_cell QUEUE operations that are only FLUSHED when the 'async with cm.get_context() as ctx:' block EXITS. Consequence for VERIFICATION: any read you do INSIDE the same block AFTER issuing an edit reflects PRE-edit state. Printing ctx.cells[name].code, inspecting globals(), or checking a returned symbol right after ctx.edit_cell(...) shows the OLD code/values — not what you just queued — because the edit hasn't been applied yet. The block's own print statements run synchronously while the queued edit is still pending. To observe the post-edit state you MUST issue a SEPARATE execute-code.sh call AFTER the mutating block has exited (its edits flush on exit, so the next block's reads see them). Workflow: (1) block A: queue edits + run_cell calls, exit; (2) block B (separate execute-code.sh): read ctx.cells[name].code / globals / run propagation checks. Do NOT try to both edit AND verify the result within one block — the read will be stale and you'll misdiagnose the edit as a no-op. Also: run_cell calls queued in the same block as edit_cell apply to the cell's NEW body only after flush, so queued runs do execute against the new code — you just can't SEE their output from within that block. Discovered build-deep-research-agent Part 3 (07-02): after ctx.edit_cell('ex1b_make_docstore', ...) the print of ctx.cells['ex1b_make_docstore'].code showed the OLD delegating body; only a follow-up block revealed the new wiring. Distinct from ctx.cells-stale-after-mutation (that's about ctx.cells being stale after DELETE/reorder; this is about edit-cell-read-back being stale within the same block).
- - DISK-FILE LAG (test-vs-disk divergence, build-deep-research-agent 07-10): code_mode edits reach the KERNEL's in-memory state when the 'async with cm.get_context() as ctx:' block exits (per above), but the DISK FILE (.py that git tracks and pytest reads) is written SEPARATELY by the marimo server — asynchronously, on browser save, or on clean shutdown. The kernel and disk are TWO different persistence layers, and the disk lags behind the kernel. SYMPTOM: after code_mode edits that create/rename/restructure cells, a pytest test reading the notebook .py file fails with 'cell X was not found' / 'def X not found in notebook file' because the disk file still has the PRE-edit cell set (or a DIFFERENT cell set from a prior kernel save). The kernel has the new cells in memory; the disk file does not — 'the notebook file on disk doesn't have startup_form or startup_validation cells at all; it has hero, intro, learning_objectives'. DIAGNOSTIC SHORTCUT: when a test fails to find a cell you KNOW you created via code_mode, grep the DISK FILE (rg 'def <cell_name>' notebooks/<file>.py) — if it returns nothing while ctx.cells shows the cell, you have disk-file lag, not a code_mode failure. FIX: (1) Stop the marimo server (clean shutdown triggers a save to disk), verify the .py file now has the cells (rg/grep), then run tests; or (2) STOP-REWRITE-RESTART (see notebook-file-editing section) — stop the server, write the correct .py on disk directly, restart. Do NOT assume code_mode edits have reached the disk file just because they succeeded in the kernel — ALWAYS verify by reading the .py file on disk BEFORE running pytest. Distinct from within-block staleness (THAT is reads within the same async-with block before kernel flush; THIS is the disk file lagging behind the kernel AFTER a successful kernel flush, sometimes by many minutes or until shutdown). Discovered build-deep-research-agent notebook 01 startup_validation/startup_form cell rename: test test_intro_notebook_has_startup_validation_cell failed because the disk file never received the code_mode rename.
## code-mode-create-cell-run-cell-id-passing
- When chaining `create_cell` -> `run_cell` in a code_mode batch, `create_cell` returns the new cell's ID into a Python VARIABLE (e.g. `hc = await ctx.create_cell(...)`). To run that freshly-created cell in the same batch, pass the VARIABLE -- `await ctx.run_cell(hc)` -- NOT a string literal of the variable's name (`run_cell("hc")`). The string `"hc"` is the NAME of your local Python variable, not a cell ID; it is neither an existing cell nor a pending-add ID, so run_cell raises 'not found in notebook or pending adds.' The asymmetry that trips you up: EXISTING cells are referenced by string literals you type ("rqfs", "FPuI"), but CREATED cells are referenced by the variable holding the returned ID. When juggling both kinds in one batch it is easy to quote a created-cell variable name by reflex. Note: run_cell DOES search pending adds (the error message says so), so passing the correct pending-add ID variable should resolve. UNRESOLVED in this session: whether queued ops (edits/creates queued BEFORE the failing run_cell) actually FLUSH when a RuntimeError fires mid-`async with` block -- the `__aexit__` receives the exception and may skip the flush, leaving the batch's apply-state uncertain; always re-read state in a SEPARATE block afterward to confirm what applied. Discovered alphaxiv-marimo-competition-submission (07-07): `run_cell("hc")` errored mid-batch after queueing edit FPuI / edit rqfs / create fi,hc,rc / move AzcX / run FPuI / run AzcX / run rqfs. Sibling of code-mode-edit-flush-on-exit (that covers read-after-edit staleness WITHIN a block; this covers referencing a just-created cell for run_cell AND the mid-block-error flush uncertainty).
- - UPDATE (07-09): the UNRESOLVED question above is now RESOLVED — see the stale-cell-error section (MID-BATCH EXCEPTION DISCARDS THE WHOLE FLUSH). Answer: a RuntimeError/StaleCellError mid-async-with block causes __aexit__ to DISCARD the flush of all earlier-queued ops; creates/edits/moves queued before the failing call do NOT apply. Mitigation: pass skip_staleness_check=True up front, or split the batch so a failure in the edit batch has no blast radius over creates.
## notebook-file-editing
- STOP-REWRITE-RESTART for major restructures (3 disk-editing escape hatches). The marimo-pair skill guard rail says "NEVER Edit/Write/NotebookEdit notebooks/*.py while a session is running" — this covers EDITS while the kernel is LIVE. There are exactly THREE legitimate escape hatches where direct file editing is the correct AND token-efficient path; in all three you STOP the server first so the file is yours: (1) Scaffolding a NEW notebook — no session yet; write on disk, then start the server. (2) Recovery from many broken/unparsable cells — dump live cell bodies via ctx.cells, STOP, rewrite notebooks/*.py with proper @app.cell function signatures, reopen. (3) MAJOR restructure of a HEALTHY notebook — rewriting the entire arc (new lesson structure, dozens of cells, full overhaul). Attempting this via dozens of edit_cell/create_cell/run_cell calls is error-prone and token-wasteful for structural changes. Instead: STOP the server, rewrite notebooks/*.py on disk, restart, reconnect. The canonical invocation (user instruction 07-07, build-deep-research-agent 6-phase arc rewrite): "stop the marimo server, rewrite the notebook, open the server again in the background, connect to it." After restart, switch BACK to code_mode (ctx.edit_cell) for surgical edits — the disk-editing exception is for the bulk rewrite only. Distinct from line-61 --watch bug workaround (that is a reactive-sync bug fix, not a deliberate restructure choice).
- CROSS-WORKTREE SESSION ISOLATION (4th direct-edit case, discovered build-deep-research-agent 07-10): the marimo-pair guard rail ('NEVER Edit notebooks/*.py while a session is running') is per-FILE-PATH, not global. A running marimo kernel managing worktree-A/notebooks/01.py will NOT clobber worktree-B/notebooks/01.py — the auto-save re-serialization writes only the file path the session opened. So when working across git worktrees (common in the build-deep-research-agent PR workflow where the marimo server runs in the PR worktree), check WHICH file path the running session actually manages before deciding direct file edits are unsafe. If the running session manages a different worktree's notebook, direct-edit the current worktree's notebooks/*.py freely — no need to stop the server or switch to code_mode. Diagnostic: 'marimo server' / discover-servers shows the session URL + the file path it opened; compare against your target file's realpath. This is distinct from the 3 STOP-REWRITE-RESTART cases (those require stopping the server because the SAME file is live); here a DIFFERENT file is live so no stop is needed.
- - PORT→WORKTREE→BRANCH VERIFICATION BEFORE EDITING VIA code_mode (discovered build-deep-research-agent 07-10): each running marimo server listens on a PORT, and that PORT maps to a specific git WORKTREE (and thus a specific BRANCH/issue). discover-servers / 'ps aux | grep marimo' / 'git worktree list' surface the mapping, but they do NOT auto-surface it to you — you must explicitly cross-reference the port against the worktree path the server process was launched from BEFORE making any edit_cell/create_cell/delete_cell call. Failure mode: you connect to port 2720 thinking it is the issue-#19 worktree, but 2720 is actually the issue-#20 (pedagogy) worktree — your code_mode edits land in the WRONG branch, silently conflating one issue's notebook changes with another. The user catches this ('hold on, 2720 is for issue #20, you're not gonna see it here'). PRE-FLIGHT CHECK (do this EVERY time before connecting to a port for the first time in a session): (1) 'git worktree list' to enumerate worktrees + their paths + branches; (2) match the port (from discover-servers or 'lsof -i :<port>') to the worktree by checking which process on that port has which cwd (ps -o cwd= -p <pid> or lsof -p <pid> | grep cwd); (3) STATE ALOUD which worktree+branch+issue the port maps to and confirm it matches the task before editing. This is the code_mode analogue of the CROSS-WORKTREE SESSION ISOLATION bullet above (that governs direct file edits; THIS governs code_mode edits — both require worktree awareness, but for different edit mechanisms). Sibling of memory on the same topic and the user-profile entry on per-issue worktree isolation.
- GIT-OPERATION CLOBBER (5th hazard, discovered build-deep-research-agent 07-10): a running marimo kernel's in-memory cell set can be STALE relative to a git merge/rebase/cherry-pick/checkout that modified notebooks/*.py on disk. The kernel does NOT re-read the file from disk on its next auto-save — it serializes its OWN in-memory graph state, silently DROPPING cells that the git operation added or changed (e.g. PR #31 merged 'startup_validation' cells; a kernel started before the merge auto-saved and wrote a file WITHOUT those cells; the agent committed the clobbered version). CRITICAL DIAGNOSTIC: 'git diff' showing NO changes is NOT proof the commit is correct — it means HEAD matches the clobbered disk version. The data loss is invisible to git because the kernel clobbered BOTH the working tree and the committed version. PREVENTION: STOP the marimo server BEFORE any 'git merge', 'git rebase', 'git checkout <branch> -- notebooks/X.py', or 'git cherry-pick' that touches a notebook file — then restart afterward so the kernel reloads from the git-updated disk. DETECTION: after committing a notebook, grep the committed file for expected cell function names (git show HEAD:notebooks/X.py | grep 'def <cell_name>') — do NOT assume 'git diff clean' means the commit captured all cells. RECOVERY: kill server, 'git checkout origin/main -- notebooks/X.py' to restore the correct version, reapply changes on disk, restart marimo, commit. This is the GIT-OPERATION sibling of main-block-clobber-on-save (that is about __main__ regeneration; THIS is about entire cells being dropped because the kernel's in-memory cell set diverged from the merged file). The existing 'direct file edits are clobbered when kernel saves' rule (marimo-pair guard rail) covers AGENT edits to the file; THIS covers GIT OPERATIONS changing the file underneath a running kernel — the same clobber mechanism, different trigger.
## reactive-graph-registration-corruption
- Symptom signature: a cell DEFINES a variable (e.g. papers = load_corpus_papers()), RETURNS it (return (papers,),), and runs WITHOUT error (no traceback in code_mode output), yet downstream cells that reference it NameError on that variable. The variable never enters marimo's shared namespace even though its defining cell executed successfully. Distinct from import-cell-return (that cell DOES return the symbol), code-mode-edit-flush-on-exit (this manifests across SEPARATE execute-code.sh blocks, not within one), and the cell erroring (there is no traceback).
Diagnostic isolation: re-run ONLY the defining cell (ctx.run_cell), then create+run a fresh MINIMAL PROBE cell (_n = len(papers),) that depends on the variable. If the probe STILL NameErrors, the defining cell's output is not registered in the reactive graph — confirming graph-registration corruption rather than a stale read or a missing return. This is the decisive test.
create_cell does NOT fix it when the CORRUPTED cell is the DEFINING cell: recreating the CONSUMING cell with a fresh cell ID (ctx.create_cell) still fails because it depends on the corrupted defining cell whose output isn't registered. The corruption lives in the defining cell's graph node, not the consumer's. You must recreate/fix the DEFINING cell — fixing only the cells that consume its output is futile.
Leading hypothesis (UNVERIFIED, build-deep-research-agent Part 3, 07-07): in-memory session-graph corruption from session restore — marimo reused an old cell ID (e.g. RGSE) from a prior session where that cell defined a DIFFERENT variable (e.g. research_store), and kept the STALE variable registration (research_store) instead of re-deriving it from the new cell body (papers). A server kill+restart did NOT clear the corruption (the same cell IDs and broken registrations persisted after reload), suggesting the graph state survives a naive restart or marimo reassigns IDs in a way that re-triggers the conflict. Needs confirmation: does the corruption clear only after a full STOP-REWRITE-RESTART with a differently-structured file?
DEFINITIVE FIX: STOP-REWRITE-RESTART (see notebook-file-editing). When edit_cell/create_cell/run_cell all fail to repair a graph-registration break, STOP trying in-session code_mode mutations — they operate on the same corrupted in-memory graph. Instead: stop the server, dump live cell bodies via ctx.cells (before stopping), rewrite notebooks/*.py on disk with clean @app.cell function signatures, restart marimo, reconnect. A freshly-loaded file rebuilds the graph from scratch. This is the recovery path for reactive-graph corruption just as it is for mass unparsable cells.
## library-code-changed-kernel-stale
- When you edit LIBRARY code (e.g. build_deep_research_agent/tools/corpus.py, solutions/part3.py — the package the notebook IMPORTS, not notebooks/*.py itself) while a marimo session is running, the kernel does NOT pick up the change. Python caches the 'from package.module import X' import on first execution, so downstream cells keep running the OLD cached module object even after you fix the .py file on disk. The marimo-pair guard rail ('NEVER Edit notebooks/*.py while a session runs') governs NOTEBOOK file edits + --watch clobbering — it does NOT apply here, but the same staleness principle does, one layer down. VERIFICATION strategy that avoids disrupting the user's active session: smoke-test the library logic directly via 'pixi run python -c "from pkg.module import fn; fn(...)"' (or 'uv run' in a uv-managed repo). This proves the on-disk library is correct independent of the marimo kernel. Do NOT restart marimo solely to verify a library fix — the notebook will pick up the new code on its next natural kernel restart, and if the notebook's cell logic is compatible with both old and new library code (e.g. it calls len() on a return value that changed from dict to dict[str,list]), it keeps running fine on the cached old code in the meantime. Note the restart need to the user rather than disrupting them mid-flow. Distinct from external-editing-stale-state (that's the --watch reactive-sync bug for notebook FILE edits) and ctx.cells-stale-after-mutation (that's ctx.cells reads after code_mode mutations) — THIS is about the kernel's PYTHON IMPORT CACHE for library modules the notebook depends on. Discovered build-deep-research-agent 07-07: fixed side_table collision in tools/corpus.py mid-session; kernel had old part3 cached; verified the fix via 'pixi run python -c "..."' instead of restarting marimo.
- IN-SESSION FIX (importlib.reload) when a cell must actually RUN the changed module code (not just verify it via the pixi-run-python-c side channel above): a verify/consumer cell that does 'import scripts.serve_X as _srv; _srv.mcp' raises AttributeError ('module has no attribute mcp') if the kernel cached the OLD module (e.g. pre-expansion file with a build_server() wrapper, before you flattened it to module-level 'mcp = FastMCP(...)'); a plain 'import' returns sys.modules' stale entry. Fix: 'import importlib; importlib.reload(_srv)' each run (or at least once after the source change). Pragmatic rule when reload-cadence reasoning gets subtle (does the kernel have old or new cached? does the cached module hold a valid docstore?): just reload BOTH the reference and starter modules EVERY run in a verify cell — verify cells run occasionally, the rebuild cost is acceptable, and it is robust against stale caches AND picks up edits. Decisive-and-correct over conditional reload logic. Discovered build-deep-research-agent Part 3 ex3_verify (07-07) after flattening serve_corpus_mcp_starter.py from build_server() to module-level mcp.
## structural-block-types
- A marimo notebook .py file is NOT just a sequence of @app.cell blocks — it can contain THREE structural block types, and a tool/agent/reviewer that walks the notebook structurally MUST account for ALL THREE or it will silently miss authored content: (1) @app.cell def <name>(...): — the reactive cells forming the visible notebook spine; (2) @app.function(hide_code=True) def <name>(...): — 'attached functions' that are HIDDEN from the default UI and often carry authored markdown specs/scaffolds (e.g. Part 4's ex1_implementation_specs holding the full 'I do' worked-example spec) — pedagogically/logically critical content a parser keyed only on @app.cell will SKIP; (3) with app.setup(hide_code=True): — the setup block holding shared imports/constants, globally visible to all cells (see unique-variable-rule). GOTCHA: a structural model that assumes 'the notebook is a sequence of @app.cell' silently drops @app.function(hide_code=True) cells — exactly where authored specs and scaffolds often live. When walking a notebook to extract prose/specs/terms, enumerate BOTH @app.cell AND @app.function (treat with app.setup(): as the shared-imports block). This is the top-level structural-completeness rule the cell-level patterns below assume but did not state. Discovered building the build-deep-research-agent tutorial pedagogy reviewer (07-07): the reviewer's structural model said 'sequence of @app.cell' but notebook 04 mixed @app.cell and @app.function, and the critical ex1_implementation_specs cell (the I-do spec) was an @app.function(hide_code=True) the reviewer would have skipped.
## hide_code-verification
- When setting hide_code=True on a cell via ctx.edit_cell(name, hide_code=True) with code=None (leaving the cell body unchanged, only toggling the config flag), the edit PERSISTS correctly — confirmed by the on-disk file showing @app.cell(hide_code=True) and the edit_cell return message 'edited code and config of cell <id>'. BUT probing the Cell view object's hide_code attribute (getattr(cell, 'hide_code', False), cell.config.hide_code, etc.) is UNRELIABLE for verification: it reads False even for cells the user MANUALLY hid via the UI AND for cells you just set hide_code=True on across a SEPARATE execute-code.sh block. Do NOT conclude 'the hide failed / did not persist' from a False probe — the probe attribute does not reflect the persisted config. Ground truth for hide_code state is the ON-DISK file (@app.cell(hide_code=True)) or the edit_cell 'edited code and config' confirmation, NOT the in-memory Cell view object. This is a narrower sibling of code-mode-edit-flush-on-exit (THAT governs ctx.cells[name].code stale reads WITHIN the same block before flush; THIS governs the hide_code ATTRIBUTE specifically, which reads False even across separate blocks and even for pre-existing user-hid cells — so it is the attribute that is unreliable, not just a same-block timing issue). Discovered build-deep-research-agent Part 1 notebook (07-07): probed hide_code on 13 narrative cells after batch edit_cell(hide_code=True) calls, all read False, triggering a misdiagnosis spiral that the hides had not taken effect — until reading the disk file confirmed all 13 had @app.cell(hide_code=True) persisted correctly and the user (who could see the UI) confirmed they were hidden.
## edit-cell-content-leaves-stale-name
- - When restructuring a marimo notebook via code_mode ctx.edit_cell (renumbering exercises, renaming sections, changing markdown headers), a CONTENT-ONLY edit changes the cell BODY but does NOT rename the cell's 'def <name>' function in the saved .py file. The old function name PERSISTS unless you explicitly pass name='new_name' to edit_cell. SYMPTOM: after editing cell content from 'Exercise 2 — Build a docstore' to 'Exercise 1 — Build a docstore', the disk file STILL shows 'def ex2_header(mo):' — the function name is now semantically stale (ex2_* implementing Exercise 1). This is a MAINTAINABILITY FOOTGUN for marimo-pair targeting: a future agent told 'edit Exercise 1's scaffold cell' searches for ctx.cells['ex1_scaffold'] and fails, because the cell is still named ex2_scaffold. The code_mode-cell-positioning-API section (above) documents that edit_cell ACCEPTS name=; THIS section warns that OMITTING it silently leaves stale names during restructures.
- ADDITIONAL DIVERGENCE: disk function names (def ex2_header) and live kernel cell IDs (auto-generated like Kclp, Hbol, emfo) can DIVERGE ENTIRELY. A notebook loaded into the kernel may assign auto-generated cell IDs unrelated to the disk function names. When debugging 'cell not found' in code_mode, check BOTH: the disk file's def names AND ctx.cells keys (which may be auto-generated IDs). Use ctx.cells to list the ACTUAL targeting keys the kernel recognizes.
- FIX WHEN RESTRUCTURING (pick by scope): (1) SINGLE CELL RENAME: pass name='new_name' to edit_cell in the same call that changes the content — ctx.edit_cell('ex2_header', code=new_body, name='ex1_header'). (2) BULK RENUMBERING (many cells): STOP-REWRITE-RESTART on disk (see notebook-file-editing section) where you control all function names directly in the .py file — attempting dozens of edit_cell+name= calls is error-prone. (3) VERIFICATION after any restructure: grep the disk file for 'def ex' and confirm every function name matches its exercise/section number — do NOT trust that content edits renamed the functions. Discovered build-deep-research-agent Part 3 notebook (03_tools_mcp_zotero.py) restructure 07-07: exercises renumbered in content via edit_cell but disk retained def ex2_header/ex2_scaffold/ex3_scaffold/ex3_try implementing Exercise 1 and Exercise 2 respectively.
## async-event-loop
- The marimo kernel ALREADY has a running asyncio event loop. Consequence: asyncio.run() and loop.run_until_complete() RAISE RuntimeError ('asyncio.run() cannot be called from a running event loop' / 'This event loop is already running') — they fail, not just behave oddly. This applies to BOTH: (1) code_mode scripts (execute-code.sh) that mutate the notebook via 'async with cm.get_context() as ctx:', AND (2) PARTICIPANT-FACING notebook cells that contain async code (e.g. a verify/smoke-test cell using 'async with Client(server) as client: ...' to test an MCP server in-process). In BOTH contexts use TOP-LEVEL 'await' + 'async with' directly — the kernel supports top-level await natively. Do NOT wrap async logic in 'async def main(): ...' + asyncio.run(main()). Secondary reason for code_mode one-liners: compound statements like 'async with' can't follow 'def name():' on the same line, so cramming into '-c' produces a SyntaxError on top of the RuntimeError. Diagnostic: if asyncio.run/loop.run_until_complete fails inside a marimo kernel (code_mode script OR a participant cell) with 'already running' / 'cannot be called from a running event loop', switch to bare top-level await/async-with — that is the fix, not a workaround to investigate. Verified build-deep-research-agent Part 3 Exercise 3 verify cell (07-07): asyncio.run failed, top-level 'async with Client(server) as c: tools = await c.list_tools()' worked directly.
## cell-id-coincidental-overlap
- Cell IDs are OPAQUE LABELS that can COINCIDENTALLY OVERLAP across sessions. After a server restart, a freshly-loaded kernel assigns IDs to cells, and an ID in the new session (e.g. 'bkHC') may MATCH the ID of a cell that was DELETED in the prior session (e.g. a removed Zotero exercise cell) — the new kernel assigned 'bkHC' to a DIFFERENT cell. This produces a confusing diagnostic state: ctx.cells lists IDs that 'look like' the old structure, tempting you to theorize 'the disk reverted / deletions were lost / IDs reassigned.' Do NOT infer cell CONTENT from a bare ID match between sessions. The DISK FILE is ground truth for content: 'rg <content> notebooks/<file>.py' settles whether a deletion persisted in ONE call (e.g. 'rg search_zotero' returning nothing proves the Zotero cell is gone even though ctx.cells still shows a 'bkHC' ID). ctx.cells IDs are authoritative for TARGETING (you must use the IDs/names the kernel recognizes to edit_cell), but NOT authoritative for CONTENT INFERENCE (the cell at 'bkHC' now is whatever the kernel loaded, which may differ from the cell at 'bkHC' before). To map a current ctx.cells entry to content, read that entry's .code or .name field — never the ID in isolation. Best practice for restart-resilient edits: target cells by NAME via ctx.edit_cell(name=...) (stable across restarts when set) rather than by auto-generated ID. Concrete instance (build-deep-research-agent notebook-3, 07-07): post-restart ctx.cells listed bkHC/PKri/Xref — the exact IDs of the DELETED Zotero cells — leading to several turns of misdiagnosis before a single 'rg' on the disk proved the Zotero content was gone and the IDs were coincidental on a fresh kernel load. Complements the existing 'ctx.cells stale after structural mutation' note (re-read after delete/reorder) and the 'disk function names vs kernel IDs diverge' note (this adds: even the SAME ID can map to different content across sessions).
## lancedb-docstore-table-name-collision
- When TWO LanceDB-backed docstores (llamabot `LanceDBDocStore` / `TurboVecDocStore`, or any LanceDB table) are constructed in the SAME marimo session/process — e.g. a notebook EXERCISE cell building a docstore with the default table_name AND a separate script/MCP server module also building a docstore with the SAME default table_name — the SECOND construction resets/corrupts the FIRST's on-disk `.lance` data files. The first docstore's in-memory handle then points at a stale/empty table, and its `retrieve`/`search` raises `lance error: Not found: .../<table-slug>.lance/data/<file>` even though the docstore object still exists in the kernel.
- DIAGNOSTIC TELL: the lance error references a table SLUG (e.g. `corpus-papers.lance`) that DIFFERS from the table your CURRENT cell/docstore uses (e.g. your cell uses `corpus_papers_mcp` → slug `corpus-papers-mcp`, but the error names `corpus-papers`). That mismatch means a DIFFERENT docstore elsewhere in the session owns the corrupted table — hunt for the OTHER docstore construction that shares the default `table_name`. The error is NOT from the cell you are running.
- ROOT-CAUSE FIX (prevention, distinct from the #440 symptom-fix of rebuilding): give each co-resident docstore a DISTINCT `table_name`. E.g. notebook Exercise 1 uses `LanceDBDocStore(table_name="corpus_papers", ...)`, while the MCP server module uses `LanceDBDocStore(table_name="corpus_papers_mcp", ...)`. This is a robustness improvement, not just a workaround — multiple docstores against the same LanceDB URI with the default name WILL collide whenever marimo re-runs their construction cells (reactive re-runs, `importlib.reload` of the script module, downstream dependents firing).
- COMMON TRIGGER in tutorial/agent notebooks: an exercise cell builds a docstore for the participant, AND a `scripts/*.py` module imported into the notebook (or a separate MCP server the notebook tests via FastMCP in-process Client) ALSO builds a docstore with the default name. When the verify cell imports/reloads the script module, the script's module-level docstore construction clobbers the exercise cell's on-disk table. Fix both sides: distinct `table_name` per logical docstore.
- DISCOVERY PATH (build-deep-research-agent Part 3, 07-07): a verify cell testing the corpus MCP server printed correct results (`mode: corpus | items: 1`) but its status transiently showed `exception` (see stale-status-after-rerun) referencing `corpus-papers.lance` (the exercise cell's table), while the verify cell itself used `corpus_papers_mcp`. The corruption came from the exercise docstore, not the verify cell — diagnosed only by noticing the slug mismatch in the error string. Resolved by giving the script's docstore `table_name="corpus_papers_mcp"`. Complements memory #440 (LanceDB stale-table error = environmental, fix by rebuilding) — #440 is the symptom/fix; THIS section is the prevention/root-cause. Reach for this when a notebook session has more than one docstore construction.
## toast-api-signature
- mo.status.toast() signature in marimo (verified 0.23.9): toast(title, description='', kind=None). Two recurring gotchas when sending toasts from code_mode or notebook cells: (1) the second positional/keyword arg is 'description', NOT 'subtitle' — passing subtitle='...' raises a TypeError (the agent's first instinct is often subtitle=); (2) 'kind' only accepts 'danger' or None (omitted) — there is NO 'success', 'info', or 'warning' kind. A success/informational toast is just mo.status.toast('Title', 'description text') with no kind= at all. When a toast call fails with an unexpected kwarg TypeError, or kind='success' is rejected, introspect inspect.signature(mo.status.toast) to confirm the installed version's exact signature — it drifts across marimo releases (same class of version-drift as the button-API-signature gotcha). Diagnostic tell: a TypeError naming an unexpected kwarg on a mo.status.* call means the installed marimo version renamed or dropped that parameter. Discovered alphaxiv-marimo-competition-submission (07-07): first toast attempt used subtitle= + kind='success' and failed; inspect.signature revealed description= + kind in {None, 'danger'}.
## progress-bar-total
- mo.status.progress_bar() used as a CONTEXT MANAGER ('with mo.status.progress_bar(...) as bar:') requires a total= argument — omitting it raises an error when the bar enters the context (it cannot infer completion without a total). Always pass total=<N> (e.g. total=len(files)), then call bar.update() inside the loop. Same version-drift / API-signature class as button-api-signature (on_click positional, label keyword-only) and toast-api-signature (description= not subtitle=, kind only 'danger'|None). When a mo.status.* call fails with an unexpected-arg or missing-arg error, introspect inspect.signature(mo.status.progress_bar) to confirm the installed marimo version's exact signature — it drifts across releases. Observed alphaxiv-marimo-competition-submission (07-07): a download cell used 'with mo.status.progress_bar() as bar:' and errored on context entry until total=len(files) was supplied.
## stale-cell-error
- ctx.edit_cell (and create_cell) raise StaleCellError when the target cell's body has CHANGED since the staleness tracker last read it. This fires ACROSS SEPARATE execute-code.sh blocks (the notebook advanced between your last read of that cell and the current edit), NOT within one block. The tracker exists to prevent clobbering concurrent edits (user edits in the browser, or marimo's own reactive re-serialization). Symptom: an edit_cell call that should work returns StaleCellError with a message like 'Read it first (e.g. ctx.cells["<id>"].code) before editing.'
- FIX (two options): (1) RE-READ FIRST — within the same get_context() block, access ctx.cells['<id>'].code (or iterate ctx.cells) to refresh the tracker, THEN call edit_cell on that same id. This is the respectful default when the user may be co-editing. (2) BYPASS — pass skip_staleness_check=True to cm.get_context(skip_staleness_check=True); safe when you know the user hasn't edited and you're iterating fast on your own queued changes.
- DISTINCTION from three look-alike staleness gotchas already in this skill: (a) code-mode-edit-flush-on-exit = stale READS within ONE block before the queued edit flushes (timing within a block); (b) ctx.cells-stale-after-mutation = stale ctx.cells after delete/reorder (structural mutation); (c) hide_code-verification = the hide_code attribute reads False even after a persisted set. StaleCellError is none of these — it is an edit_cell REFUSAL across blocks because the cell drifted from what you last saw. Discovered alphaxiv-marimo-competition-submission (07-07): edit_cell on cell AzcX after it had errored (and been modified by the error/re-run cycle) raised StaleCellError; resolved by re-reading ctx.cells['AzcX'].code in the new context, then editing.
- - MID-BATCH EXCEPTION DISCARDS THE WHOLE FLUSH (RESOLVED): when an exception (StaleCellError on a later edit_cell, OR a RuntimeError, OR any synchronous validation error) fires inside the 'async with cm.get_context() as ctx:' block, the __aexit__ does NOT flush the earlier-queued ops — they are DISCARDED. A batch that does [create_cell x5, run_cell x2, edit_cell(Vxnm)] where the edit_cell raises StaleCellError results in ZERO cells created: the 5 create_cells were queued synchronously and SHOULD have flushed on block exit, but the propagating exception aborts the flush, so nothing lands. This is non-obvious because Python's normal 'async with __aexit__ runs even on exception' intuition suggests the flush should still happen — it does not. Confirmed by verification: after the exception, ctx.cells showed NONE of the queued create_cells existed. This RESOLVES the previously-UNRESOLVED question in the code-mode-create-cell-run-cell-id-passing section below. MITIGATION (pick one): (1) PREFERRED — pass skip_staleness_check=True UP FRONT when you intend to batch create + edit together: 'async with cm.get_context(skip_staleness_check=True) as ctx:'. This removes the most common mid-batch exception source (the edit hitting staleness) so the whole batch flushes cleanly. Safe when you know the user isn't co-editing. (2) SPLIT THE BATCH — do create_cells in one block (flush on exit), then edit_cell in a SEPARATE block, so a failure in the edit batch has no blast radius over the creates. (3) Always re-verify in a SEPARATE execute-code.sh block afterward — never assume queued ops landed after an exception; re-read ctx.cells to confirm what applied. The cost of skipping the mitigation: ~2 reasoning rounds + a verification call to rediscover that 'queued != applied on exception' — this session burned that. Discovered build-deep-research-agent Part 3 stretch-exercises creation (07-09): edit_cell('Vxnm', recap) raised StaleCellError after 5 create_cells + 2 run_cells were queued; re-check showed zero cells created; redo with skip_staleness_check=True landed all 5 cells + the recap edit cleanly.
## description
- Common pitfalls when writing and editing marimo notebook .py files. Covers: import patterns (never alias marimo as mo), cell parameter requirements (mo must be received as parameter), app initialization (use marimo.App() not mo.App()), expression capture in try/except blocks, unique variable rules, import cell return values (imported symbols must be returned for cross-cell propagation), button API signature (on_click is first positional arg, label is keyword-only), toast API signature (mo.status.toast uses description= not subtitle=, kind only "danger" or None, no "success"), StaleCellError on edit_cell (re-read ctx.cells or use skip_staleness_check=True), progress_bar requires total= as a context manager, UNDERSCORE-PREFIX PRIVACY (names starting with _ are cell-local and do NOT propagate to other cells via the reactive graph — rename to non-underscore to share across cells), and mo.ui.slider API (uses start/stop/step and a numeric Sequence 'steps' param for discrete snapping — there is NO 'options=' kwarg; use mo.ui.dropdown with a {label: value} dict for LABELED discrete choice).
- - Addendum to description: also covers mid-batch-exception-discards-flush (a StaleCellError or any synchronous exception raised inside an 'async with cm.get_context()' block causes __aexit__ to DISCARD the flush of all earlier-queued create_cell/edit_cell/run_cell ops — nothing lands; mitigation = pass skip_staleness_check=True up front when batching create+edit, or split the batch so an edit-batch failure has no blast radius over creates).
- Addendum: also covers the cell.errors diagnostic attribute (list[CellError] on each cell object — complementary to cell.output.data), and the stale-status paradox where c.errors returns empty (0 errors) while cell.status reads exception (stale from a prior run; re-run to clear).
## underscore-prefix-cell-local
- In marimo, a variable name starting with an UNDERSCORE (e.g. `_vis_model`, `_L0`, `_df`) is NOT cell-local and NOT private in the Python-underscore sense. marimo STRIPS a single leading underscore from a cell definition when injecting it into the shared reactive namespace: a cell that defines `_df` actually shares it with descendant cells under the name `df` (underscore removed). The value DOES enter the reactive graph and DOES propagate — just under the stripped name. So a descendant CANNOT see `_df`, but CAN see `df`. This stripping is why underscore names APPEAR 'not to propagate' — they do propagate, under the stripped name. CANONICAL MECHANISM + diagnostic signature live in marimo-pair/reference/gotchas.md (#private-variables-are-cell-scoped-underscore-stripping-mechanism); this section is a pointer, not the source of truth. SYMPTOM: cell B references `_vis_model` defined in cell A and gets NameError ('_vis_model' is not defined) with the SMOKING-GUN hint `Did you mean: '<name-without-underscore>'?` (e.g. `Did you mean: 'anim_graph_choice'?`), even though cell A ran cleanly and you can see `_vis_model` in cell A's own output. FIX: rename the name to NON-underscore (`_vis_model` -> `vis_model`, `_L0` -> `L0` or `vis_L0`) so marimo registers it as a reactive variable that propagates across cells. Keep underscore-prefixed names ONLY for genuinely cell-local intermediates (loop counters, throwaway transforms used once within the same cell) — anything another cell needs must drop the leading underscore. RULE OF THUMB when authoring a marimo cell: ask 'does any OTHER cell reference this name?' — if yes, no leading underscore. This is a BEHAVIORAL rule of marimo's reactive graph, distinct from the user's general stylistic preference for non-underscore Python names (memory #196): even a user who likes underscore-prefixed 'private' helpers in plain Python MUST drop the underscore in marimo for any cross-cell value, or the value simply will not propagate. NOTE: this interacts with unique-variable-rule — every cross-cell (non-underscore) name must still be defined in exactly ONE cell, so when you un-underscore a name to share it, make sure no other cell already defines that bare name (e.g. `vis_model` is safe only if no other cell defines `vis_model`). Discovered alphaxiv-marimo-competition-submission (07-07): the heatmap cell defined `_vis_model`, `_vopt`, `_perm`, `_idx`, `_lo`, `_L0`, `_gamma` — all correctly private to the heatmap cell EXCEPT that `explore_view` in a different cell needed `_vis_model[0].weight`, which raised NameError; renaming `_vis_model` -> `vis_model` (leaving the other _ vars private) propagated it and the explorer resolved.
CORRECTION (07-18, Network-Analysis-Made-Simple): an earlier version of this bullet claimed underscore-prefixed names are 'PRIVATE / cell-local and do NOT enter the reactive dataflow graph' — that mechanism is WRONG (it is underscore-STRIPPING on share, not graph exclusion). The wrong claim caused an agent to spend many turns theorizing about dependency-graph wiring and staleness for a NameError on `_anim_graph_choice` whose error message literally said `Did you mean: 'anim_graph_choice'?` — the stripping signature. When you see that `Did you mean:` hint on a leading-underscore name, the underscore IS the cause; rename the definition to non-underscore. The user's emphatic correction 'no _private_variables!!!!!' (5 exclamation marks, the 3rd time stating this rule: 02-22, 07-10, 07-18) was enforcing the standing no-underscore naming preference (memory #618) — do NOT reinterpret an emphatic repetition of a known rule as a novel technical claim about marimo internals.
## ui-slider-dropdown-api
- mo.ui.slider API in marimo: the constructor uses POSITIONAL/RANGE params — slider(start, stop, step=None, value=None, ...) for a continuous or fixed-step range, and an additional `steps` kwarg accepting a Sequence[Numeric] (e.g. steps=[2,3,4,8,16,32]) for DISCRETE snapping to a custom set of values. There is NO `options=` kwarg on mo.ui.slider — passing options= raises TypeError. This is the same version-drift / API-signature class as button-api-signature (on_click positional, label keyword-only), toast-api-signature (description= not subtitle=), and progress_bar-total. TWO CHOICES for a discrete selector, with tradeoffs: (1) mo.ui.slider(steps=[2,3,4,8,16,32], value=3) — gives a TACTILE DRAG through the discrete values, but the slider DISPLAYS the raw numeric value (e.g. '3') with no label, so you must interpret/annotate it in the view (markdown like 'lower = fewer bits'). (2) mo.ui.dropdown(options={'1.58-bit ternary': 3, 'full precision (32-bit)': 32, ...}) — gives LABELED discrete choice (dict maps {label: value}, and .value returns the dict VALUE not the label), but requires a CLICK to change (less 'tactile' than a drag). Decision heuristic: if the value IS self-describing as a number and you want drag feel (precision knob, level count), use slider+steps; if you need human-readable labels per option (named quantization schemes, named configs), use dropdown. For labeled options that should STILL feel interactive, consider mo.ui.radio_group or mo.ui.dropdown with `orientation='horizontal'`. When ANY mo.ui.* call fails with an unexpected-kwarg TypeError, run inspect.signature(mo.ui.slider) / inspect.signature(mo.ui.dropdown) to confirm the installed marimo version's exact signature — the UI widget API drifts across releases. Discovery path that cost multiple probing turns (alphaxiv-marimo-competition-submission, 07-07): assumed slider(options=...) -> TypeError -> discovered start/stop/step -> discovered the `steps` Sequence param -> weighed slider-with-steps vs dropdown-with-dict -> landed on slider(steps=[2,3,4,8,16,32], value=3) for a 'drag the precision' BitNet quantization explorer.
## loop-variable-cell-collision
- LOOP VARIABLES (for-loop targets) become PUBLIC marimo names and collide across cells, exactly like 'fig'/'ax'/'mo'/'Counter' do. A 'for i in range(...)' in cell A makes 'i' a reactive variable owned by cell A; a 'for b, m, l in zip(...)' in cell B makes 'b', 'm', 'l' public too. When a THIRD cell uses 'for i, m in enumerate(...)', marimo raises a multiple-definition / batch-rollback validation error because BOTH 'i' and 'm' are already owned by other cells. This is the LOOP-VARIABLE analog of the matplotlib 'fig, ax' collision (already documented) and the import-alias collision, but EASIER to hit because (1) loop vars like 'i','j','k','m','n','b' are common throwaway names every cell reuses, and (2) the loop variable LEAKS into the cell's public namespace even though it 'feels' local. FIX: namespace loop variables per-cell with a short unique prefix derived from the cell's purpose — e.g. a heatmap cell uses 'hi, hj'; a frontier_plot cell uses 'fi, fm'; an efficiency cell uses 'ei, em'. Do NOT rely on Python scoping to localize loop vars — marimo's reactive graph treats every for-loop target as a cell output. PRE-FLIGHT when adding ANY new cell that contains a 'for' loop to an EXISTING notebook: grep all cells for 'for <var>' / '<var> =' to see which common loop-name letters ('i','j','k','m','n','b','r','x','y') are ALREADY owned, then pick an unused prefix for the new cell's loop vars. This avoids the silent batch-rollback (the entire create_cell batch is rejected, so NONE of the new cells are created — you must recreate the whole batch with the renamed vars, wasting a round-trip). Concrete instance (alphaxiv-marimo-competition-submission, 07-07): a batch of 4 new cells was rolled back because eff_plot's 'for i, m in enumerate(eff_mems)' collided with 'i' (heatmap's 'for i in range(0, len(x_train), 1024)') and 'm' (frontier_plot's 'for b, m, l in zip(...)'); fix was renaming to 'for ei, em in enumerate(eff_mems)' and recreating all 4 cells. Sibling of the matplotlib 'fig, ax' collision and the unique-variable-rule — extends the unique-variable constraint to the FOR-LOOP TARGET class of assignment, which the existing sections cover only for plt result objects and imports.
## remote-molab-sandbox-extraction
- ## remote-molab-sandbox-extraction
When working with a PRIVATE molab sandbox (sb-XXXXX.sb.molab.run) via marimo-pair's execute-code.sh --url, the kernel runs on the REMOTE Modal server, NOT localhost. Three non-obvious gotchas (alphaxiv-marimo-competition-submission, 07-07):
1. **molab sandbox URLs do NOT serve .py via Accept header.** Unlike PUBLIC molab.marimo.io /notebooks/nb_<id> links (which return source via curl -H 'Accept: text/x-python'), a private sandbox URL returns an HTML page SHELL. Do NOT waste a round trying curl+Accept on a sandbox URL.
2. **execute-code.sh output capture fragility.** Piping execute-code.sh output through 'sed 1d' with heredoc + '2>/dev/null' ('bash execute-code.sh <<EOF 2>/dev/null | sed 1d > file') produced a 0-BYTE file — the heredoc+pipe+stderr-redirect combination silently fails to capture stdout. Workaround: redirect RAW stdout+stderr to a file first ('bash execute-code.sh <<EOF > file 2>&1'), inspect the file, then clean (strip the 'Warning: connecting...' line, HTML wrappers like <pre>) in a second pass. Do NOT pipe through sed on the first attempt.
3. **Subagents need LOCAL files.** A review subagent (Task tool) runs on the local machine and CANNOT access the remote molab kernel or remote /marimo/notebook.py. To review a remote notebook's content: (a) extract cell codes to a LOCAL file via execute-code.sh (print each cell's name+code, capture raw stdout to a local .md/.txt — a valid .py is NOT needed for review, just codes in order); (b) dispatch the subagent to Read the local file; (c) apply fixes back through ctx.edit_cell.
**marimo-pair skill prose correction:** the marimo-pair SKILL.md 'Verifying Rendered Output' section states 'The kernel runs on the local host, so the file lands on the machine Read can see.' This is FALSE for molab sandboxes — a file written to /tmp/ inside the kernel lands on the REMOTE box. For molab, the only reliable visual-QA path is to ASK THE USER (who sees the notebook live in their browser). Already noted in memory #466 (screenshot/visual-QA limitation) but the marimo-pair prose was never corrected. Sibling of memory #460 (public molab link source fetching) and #466 (remote-kernel visual-QA limitation).
## ctx-cells-iteration-yields-objects
- When iterating marimo-pair code_mode's ctx.cells (a _CellsView), 'for cid in ctx.cells' yields cell OBJECTS (NotebookCell instances), NOT string IDs — it does NOT follow Python's mapping convention where __iter__ yields keys. Passing the iterated value as a lookup key to ctx.cells[cid] fails with: 'Cell NotebookCell(id=...) not found. Available cell IDs: [PZww, ...]' — the error proves the iterated cid is a NotebookCell object, not a string. CONFIRMED: list(ctx.cells.keys()) returns string IDs; ctx.cells.items() yields (id_string, cell_object) tuples and works correctly. FIX: use 'for cid, c in ctx.cells.items()' to get (id, cell) pairs directly (no lookup needed — you have c already), OR use 'for cid in ctx.cells.keys()' for explicit string-ID iteration if you need the lookup. Do NOT write 'for cid in ctx.cells:' then 'ctx.cells[cid]' — the cid is a cell object and the lookup fails. This is the ITERATION sibling of ctx.cells-stale-after-mutation (stale reads after delete/reorder), cell-id-coincidental-overlap (IDs reused across sessions), code-mode-edit-flush-on-exit (stale reads within one block), and stale-cell-error (edit_cell refusal across blocks) — all four masquerade as 'cell not found' but have distinct root causes; THIS one is a wrong-iteration-protocol bug, not staleness. Discovered alphaxiv-marimo-competition-submission (07-07): agent wrote 'for cid in ctx.cells: print(ctx.cells[cid].code)' and hit 'Cell NotebookCell(...) not found', reasoned through _CellsView.__iter__ semantics for several steps, landed on .items() as the robust fix.
- ADDITION (07-10): _CellsView is NOT a full dict — it supports [] key-indexing, .keys(), .values(), .items(), but does NOT support .get(). An agent calling ctx.cells.get('name') hits AttributeError ('_CellsView' object has no attribute 'get'). FIX: build a name->cell dict first ('cells_by_name = dict(ctx.cells.items())' then 'cells_by_name.get(...)'), OR access with ctx.cells['name'] inside try/except KeyError, OR (preferred) just iterate ctx.cells.values()/items() and filter by cell.name. General principle: treat _CellsView as a read-only mapping with a LIMITED API (no .get/.pop/.setdefault), not a dict subclass — do not assume dict methods exist. This is the .get() sibling of the __iter__-yields-objects gotcha above: both stem from _CellsView not following the dict protocol. Discovered build-deep-research-agent 07-10: agent wrote ctx.cells.get(nm) and hit the AttributeError, burned a turn reasoning about _CellsView before rebuilding a dict from .items().
## wigglystuff-library
- wigglystuff (koaning / Vincent D. Warmerdam, v0.5.13+, anywidget-based) is the go-to INTERACTIVE-WIDGET library for marimo notebooks — esp. for 'guided walkthrough of the notebook' requests. Do NOT guess its widget set from memory: common wrong guesses (Tag, Checkbox, Radio, BouncingLogo, PieChart, Scatter, ImagePair, GithubCard, Spinner, Confetti) are NOT in wigglystuff. Either introspect 'import wigglystuff as ws; [n for n in dir(ws) if not n.startswith("_")]' after install, or use this VERIFIED inventory (07-07, 0.5.13). Headline widgets: CellTour + DriverTour (guided tours that highlight marimo cells or DOM elements in sequence — EXACTLY what a 'guided walkthrough' means), TangleSlider (inline Bret-Victor-Tangle slider), TangleChoice / TangleSelect (inline toggle/choice), Matrix (editable or static 2D matrix display — great for weight matrices), PlaySlider (slider with play button), ProgressBar, SortableList, Slider2D, CircularSlider / CircularRangeSlider. Drawing/creative: Excalidraw, Paint, BezierCurve, SplineDraw, CurveEditor, EdgeDraw, GridDraw, Path. Data-viz: AltairWidget, ParallelCoordinates, RidgelineChart, ScatterWidget, Treemap, NestedTable, GraphWidget, Neo4jWidget, WandbChart, ChartMultiSelect / ChartPuck / ChartSelect (chart-picker widgets). Utility: CopyToClipboard, ColorPicker, HoverZoom, KeystrokeWidget, WebcamCapture, WebkitSpeechToTextWidget, GamepadWidget, LiveEdit, TextCompare, AnnotationWidget, EnvConfig, ApiDoc, ModuleTreeWidget, HTMLRefreshWidget, ImageRefreshWidget, ThreeWidget, forecast_chart.
CELLTOUR STEP SCHEMA (the headline recipe for a notebook walkthrough): ws.CellTour(steps=[{"cell": <int-index>, "title": "...", "description": "..."}, ...], auto_start=True, show_progress=True). Steps reference marimo cells by INTEGER INDEX (0-based, in notebook order) OR by the cell's NAME via a SEPARATE 'cell_name' key (e.g. '{"cell_name": "hero"}'). IMPORTANT: 'cell' and 'cell_name' are TWO DISTINCT KEYS in the step dict (verified against wigglystuff cell_tour.py _transform_step 07-08): 'cell' -> integer INDEX (used as the nth .marimo-cell); 'cell_name' -> name STRING (compiled to a [data-cell-name="..."] CSS selector). Do NOT write '{"cell": "hero"}' for a name-based step — that passes the string 'hero' as the integer-index slot and silently fails to resolve. Name-based is '{"cell_name": "..."}'. Verify the current cell count/order before building a tour that uses indices — indices shift when cells are added/reordered (a cell named via @app.cell decorator's name attribute, or a cell you named in marimo-pair, is the more stable reference). DriverTour is the DOM-element analog (CSS selectors) for touring non-cell page elements.
INSTALL: anywidget lib — conda/pypi MIXED solve conflicts in pixi projects, so add as a pypi-dependency: 'pixi add --pypi wigglystuff' (resolves to '>=0.5.13,<0.6'). In a marimo-pair/running-kernel session use 'ctx.packages.add("wigglystuff")' (installs cleanly, usually NO kernel restart needed). Docs: koaning.github.io/wigglystuff. PyPI rendered page is JS-gated/empty — use the PyPI JSON API (pypi.org/pypi/wigglystuff/json) for metadata (see pypi-package-metadata-json-api skill).
DISCOVERY ANTI-PATTERN: when a user asks for a 'guided walkthrough' / 'tour' / 'interactive walkthrough' of a marimo notebook, the answer is ws.CellTour (marimo cells) or ws.DriverTour (DOM) — do not hand-roll a stepper from mo.ui.tabs + mo.ui.radio. Recurring (alphaxiv-marimo-competition-submission 07-07 hero→close cell tour; build-deep-research-agent notebooks/03_tools 07-07 'talking tool' walkthrough). For inline-Tangle-style interactivity (e.g. an interactive 'feel the 1.58-bit knee' slider), ws.TangleSlider; for a matrix viz (e.g. ternary {-1,0,1} weight matrix), ws.Matrix.
EDIT-SEQUENCE PROCEDURE (build the CellTour LAST, verify refs before 'done'): in a multi-phase marimo-pair edit session on a notebook that has (or will have) a CellTour, order the work as (1) STRUCTURAL moves — delete/move/create cells, (2) CONTENT rewrites — edit_cell code/markdown, (3) REBUILD the CellTour in the FINAL phase in a FRESH context, then (4) VERIFY every tour step's cell reference resolves to a live cell before declaring done. This defends against three failure modes that each silently break the tour with 'missing refs': (a) a tour rebuilt MID-context reads STALE cell positions because code_mode ops queue and flush on context exit (see code-mode-edit-flush-on-exit) — always rebuild in a fresh context AFTER all moves have flushed; (b) a named cell's name can silently reset to '_' after an edit_cell (see the name-vanish-to-underscore rule below), so a name-based step '{"cell_name": "fix"}' stops resolving; (c) index-based steps shift on any reorder/delete. Verification step: after rebuilding, iterate the tour's cell values and confirm each maps to an existing cell — for name-based steps grep the disk 'def <name>(' or re-read ctx.cells[id].name; for index-based steps re-derive indices from a FRESH ctx.cells read (never from memory of an earlier read) and confirm index < current cell count AND points at the intended cell. Prefer NAME-based steps over index-based — names survive reorders, indices don't. Recurring (alphaxiv-marimo-competition-submission 07-08: user 'make sure the celltour stuff is redone correctly, AFTER fixes to the notebook are in place. we had missing refs the last time round we edited, so the celtour broke').
## plotly-over-matplotlib
- For marimo notebook charts, prefer Plotly (go.Figure / plotly.express) over matplotlib. Plotly figures render NATIVELY and interactively in marimo (hover, zoom, pan) — just `return fig` from the cell — whereas matplotlib is static. Stated 07-07 (alphaxiv-marimo-competition-submission): 'we should be using plotly instead of matplotlib here.' This is the marimo-notebook sibling of the user-profile blog-post rule (#25: Plotly.js CDN-embedded in Lektor bodies) — same interactive-over-static preference, different surface.
- CONVERSION RECIPE (matplotlib cell → plotly cell): `plt.subplots()` → `go.Figure()`; `ax.scatter/ax.plot` → `go.Scatter(mode='lines'|'lines+markers', error_y=...)` ; `ax.bar` → `go.Bar`; `ax.imshow` → `go.Heatmap`; a matplotlib `LinearSegmentedColormap` (e.g. a diverging cmap from indigo→neutral→amber) → a plotly `colorscale` list of `[pos, color]` tuples. Styling: `fig.update_layout(height=, paper_bgcolor='white', plot_bgcolor='white', margin=, font=, title=, legend=, showlegend=)`; `fig.update_xaxes/update_yaxes(title_text=, gridcolor='#eef2f7')`. Keep per-cell UNIQUE figure variable names (rfig, effig, lmfig, efig...) — never reuse bare `fig`/`ax` across cells (see the 'matplotlib fig,ax plt.subplots collision' rule; the collision rule names matplotlib but the unique-name rule applies to plotly too).
- REMOVING matplotlib is a REACTIVE CASCADE: if you delete `import matplotlib` (and the cmap) from the imports cell, marimo reactively re-runs every cell that still uses plt → they ALL error. Convert ALL still-matplotlib cells in one sweep (the cells will error until each is converted), do NOT convert one cell at a time and report 'clean' mid-sweep. Discovered 07-07 converting 6 cells (heatmap, frontier, race, eff, explore, lm) — removing matplotlib from imports cascaded the 3 unconverted cells into errors until Batch 2 converted them.
## live-chart-during-synchronous-loop
- A single marimo cell runs its body SYNCHRONOUSLY — it cannot yield to marimo's reactivity mid-execution. So you CANNOT update a returned chart object inside a `for step in range(...)` training loop and have marimo re-render it each iteration: the cell only returns once, at the end. `mo.status.progress_bar` is the only built-in loop-friendly primitive, and it shows a 0-100% bar, NOT the live trajectory. When the user wants the LIVE TRAINING CURVE (loss vs step updating in real time, NOT a progress bar — stated 07-07), the synchronous-loop constraint forces a widget-based approach.
- PATTERN THAT WORKS (live curve during a tight loop): an `anywidget.AnyWidget` subclass holding a traitlet list of `[step, loss]` points, whose `_esm` renders an SVG/canvas line chart and redraws on `model.on('change:points')`. During the Python training loop, every K steps append points (`chart.points = chart.points + [[step, loss]]`) — the traitlet syncs to the frontend and the JS redraws WITHOUT full cell re-execution or figure re-serialization. This is the marimo-native live-chart mechanism; a 'rebuild and return a fresh plotly fig each iteration' approach only works if the loop is slow enough to tolerate full figure re-serialization per step.
- Distinct from wigglystuff's `forecast_chart`/`HTMLRefreshWidget` (which refresh on a timer/interval, not on loop-driven traitlet pushes) and from `mo.ui.refreshable` (re-runs the whole cell, not a mid-loop update). The anywidget-traitlet approach is the right tool when you need the chart to advance INSIDE a single synchronous training loop. Pair with the plotly preference: for a NON-live (final) chart use plotly go.Figure; for a LIVE-during-loop chart use the anywidget SVG approach (plotly re-serialized per step is too heavy for a tight training loop).
- RENDERING FAILURE MODE (data present but chart blank): the anywidget pattern above can FAIL TO RENDER in marimo even when the Python-side widget object has correct data (verified: widget.series / widget.points has all the expected entries, e.g. 50 (step,loss) points after training). The user sees NOTHING — no chart, just empty space or a stale 'warming up' placeholder. This is a RENDERING failure, not a data failure. THREE candidate root causes: (1) CELL SPLIT — the widget is DEFINED+DISPLAYED in cell A (e.g. live_chart) and its series/points are MUTATED by a training loop in cell B (e.g. live_train). This conflicts with the anywidget-single-cell rule earlier in this skill (stateful anywidget components should live in ONE cell). When marimo's reactive graph runs all cells, cell B mutates the widget during its synchronous execution; if cell A re-runs afterward (reactive edge) the widget is recreated empty. Diagnostic: does the widget object STILL have data after the notebook settles? If yes, cell A did not re-run and the problem is the frontend, not the data flow. (2) COMM TIMING — the traitlet 'change:points'/'change:series' event fires DURING cell B's synchronous execution, potentially BEFORE the widget's _esm has attached its SVG/canvas to the DOM (marimo batches display updates). The draw() function runs against a not-yet-attached element and the update is silently lost; no subsequent change event re-triggers it. (3) NO INITIAL-RENDER DRAW — the _esm registers draw() ONLY on model.on('change:points', draw) but does NOT call draw() on INITIAL render. If the data is already populated by the time the widget is displayed (data set during the prior cell's execution, display batched later), draw never fires and the canvas stays empty. DIAGNOSTIC METHODOLOGY (verified useful): FIRST verify Python-side data (does widget.series have the expected entries? if yes, STOP debugging data flow — the problem is rendering); THEN inspect the _esm to check whether draw() fires on BOTH initial render AND change events; THEN check whether the widget is split across cells. FIX DIRECTIONS (the session that discovered this was cut off before confirming which fix resolved it): (a) ensure the _esm calls draw() on INITIAL render in addition to on change:series — add an explicit draw() call after the SVG is appended in the render function, so pre-populated data renders on display; (b) consolidate widget DEFINITION + training LOOP into a single cell so there is no cross-cell mutation timing gap; (c) verify _css is not scoped out by marimo's anywidget integration (the .llc-svg{height:...} rule may need to be inline or in a <style> inside the _esm rather than in _css). Discovered alphaxiv-marimo-competition-submission 07-07 (caveats the PATTERN THAT WORKS bullet above which recommends the exact approach that failed).
- CONFIRMED 2-CELL FAILURE (alphaxiv-marimo-competition-submission 07-07): the two-cell anywidget-traitlet pattern — widget DEFINED+DISPLAYED in cell A (live_chart) and its .series/.points MUTATED inside cell B's training loop (live_train) via 'live_chart.series = [{...,points:pts}]' + 'await asyncio.sleep(0)' — renders ONLY THE FIRST UPDATE to the frontend. User-confirmed symptom: after running live_train, the chart showed legend + axes but y-axis 2.30-3.30, x-axis 'step 1', and NO line — exactly one point (step 0, loss 2.30) reached the browser. Subsequent mutations queued during the blocking loop did NOT flush, and bumping the yield to 'await asyncio.sleep(0.03)' (30ms) did not resolve it. This CONFIRMS root cause #2 (COMM TIMING) above: 'asyncio.sleep' yields the kernel event loop but does NOT reliably flush marimo's widget-comm channel to the frontend during a synchronous cell body — the first update flushes (initial render) but later ones are silently lost. So the anywidget-traitlet approach only works reliably in a SINGLE cell, not across two; the earlier bullet that described the cross-cell mutation as if it were the working pattern is INCOMPLETE — cross-cell widget mutation during a blocking loop does not propagate.
CONFIRMED ROBUST FIX (VERIFIED 07-08 — the alphaxiv-marimo-competition-submission notebook shipped this exact pattern in cells/jit_train.py lines 13 & 39 in a 0-error submission): merge the chart widget and the training loop into ONE cell and use 'mo.output.replace(widget)' to push live updates. 'mo.output' (marimo's mid-cell output API: 'append', 'replace', 'replace_at_index', 'clear') replaces the cell's own output area mid-execution, so a single cell can: (1) create the anywidget or plotly chart, (2) 'mo.output.replace(chart)' to mount it, (3) inside the training loop, mutate the chart then 'mo.output.replace(chart)' every K steps — each replace explicitly commits the output to the frontend without relying on a traitlet-comm change event flushing. This sidesteps BOTH the cell-split problem (root cause #1) and the comm-timing problem (root cause #2) because the output is committed each iteration rather than pushed through a change:points event. VERIFIED PATTERN (jit_train.py): mount the figure with mo.output.replace(curve_fig) BEFORE the loop, then inside the loop add trace points and call mo.output.replace(curve_fig) every K steps (the submitted notebook used step % 200). The every-K-steps cadence confirms this produces INCREMENTAL live updates (a new point appears on the running curve), not a single end-of-loop render. Fallback if mo.output.replace is ever batched in a future marimo version: break the training into per-step cells driven by mo.ui.run_button, OR write the loss series to a file and poll it from a separate refreshable cell.
- mo.output.replace is a mid-cell output-commit API, distinct from traitlet-comm change events (change:points) which fire asynchronously and may be batched/dropped during a synchronous loop. The hypothesis motivating mo.output.replace over the traitlet approach: an explicit output-commit call forces marimo to push the rendered widget to the frontend at that point in the loop, whereas a traitlet mutation relies on marimo's comm channel flushing (which the CONFIRMED 2-CELL FAILURE shows does not happen reliably mid-loop).
## scratch-cell-cuda-poisons-shared-kernel
- - A FATAL CUDA error (device-side assert) raised in a SCRATCH / prototype cell POISONS the shared marimo kernel's CUDA context for EVERY GPU cell in the live notebook — not just the cell that crashed. Root cause: all marimo cells (participant cells AND code_mode execute-code.sh scripts AND ad-hoc scratch cells) run in the SAME Python process / SAME kernel, so they share ONE CUDA context. A PyTorch CUDA device-side assert (e.g. passing an input sequence longer than a transformer's max_position_embeddings — see transformer-context-length-cuda-assert skill) corrupts that single context irrecoverably; every subsequent torch GPU op then fails with a misleading secondary error until the kernel is RESTARTED. SYMPTOM: your notebook's lm_experiment / training cells were running fine, you prototyped something in a scratch cell that triggered a CUDA assert, and now re-running those previously-fine cells ALSO fails with CUDA errors — the scratch assert poisoned them. There is no in-process recovery and NO reliable code_mode API to restart the marimo kernel (ctx.execute_command / enqueue_command do not expose a documented 'restart kernel' command), so you MUST ask the user to restart the kernel in the marimo UI (the notebook autosaves, so restart + Run All rebuilds everything). PREVENTION: never eval/fine-tune a pretrained model on unbounded text in a scratch cell — always tokenize with truncation=True, max_length=<model.config.max_position_embeddings> (1024 for GPT-2/distilgpt2), or chunk long corpora into <=max_length windows. Prototype risky GPU ops in an ISOLATED process (pixi run python -c '...' / uv run) rather than in the shared marimo kernel when the cost of poisoning the live notebook is high. Distinct from library-code-changed-kernel-stale (stale Python import cache — recoverable via importlib.reload, no kernel restart needed) and async-event-loop (RuntimeError, recoverable by switching to top-level await): a CUDA device-side assert is NON-recoverable and MANDATES a kernel restart. Discovered alphaxiv-marimo-competition-submission (07-07): a QAT/STE fine-tuning-recovery prototype eval on a 4000-char (~9000-token) text exceeded distilgpt2's 1024 context, hit a fatal index-out-of-bounds assert in the positional-embedding lookup, and broke the live notebook's other GPU cells until the user restarted the kernel.
## underscore-convention-vs-user-preference
- TENSION with user style: this skill documents underscore-prefix cell-local privacy ('_names do NOT propagate' in marimo — factually correct marimo behavior), but the user STRONGLY DISLIKES leading-underscore names everywhere (user-profile #33: 'I really don't like _private_methods!!!! get rid of them'). When authoring marimo notebooks FOR THIS USER, do NOT use the underscore-prefix technique even though marimo supports it — the no-underscore preference overrides it. The challenge: marimo ENFORCES unique variable names across ALL cells, and underscore-privacy was the cheap way to dodge that constraint. RESOLUTION that satisfies BOTH constraints: give every cell a DESCRIPTIVE PREFIX and namespace loop variables per cell WITHOUT underscores — e.g. a 'fix' cell uses fix_/ei/em; a 'live_train' cell uses live_/li/lm; a 'heatmap' cell uses heat_/hi/hj. This keeps every top-level name globally-unique (satisfying marimo) AND public/descriptive (satisfying the user). Grep 'for <var>' across the notebook before adding a loop cell to pre-empt collision. Discovered alphaxiv-marimo-competition-submission 2026-07-07: a 29-cell notebook had ~all names underscore-prefixed (_idx, _RoundSTE, _fix_qwt, _HFC, etc.) and required a full refactor to per-cell public prefixed names.
- TWO EXTENSIONS discovered alphaxiv-marimo-competition-submission 2026-07-07 (same 29-cell notebook): (1) IMPORT-ALIAS COLLISION for SHARED LIBRARIES needed at CLASS-DEFINITION TIME — when two cells both need the SAME library (e.g. anywidget, traitlets) at top-level because each DEFINES a class subclassing it (e.g. 'class PrecisionFader(anywidget.AnyWidget)' using 'traitlets.Int'), reactive propagation does NOT help (the class body needs the module at DEFINITION time, not at cell-run time via the reactive graph) and you cannot do 'import anywidget' in both cells (MultipleDefinitionError on 'anywidget'). The original code dodged this with UNDERSCORE aliases ('import anywidget as _anyw' / 'import traitlets as _trt' — cell-local, no collision) but that violates the no-underscore preference. FIX: alias the IMPORT STATEMENT ITSELF with a descriptive NON-UNDERSCORE name per cell — 'import anywidget as explore_anywidget' / 'import traitlets as explore_traitlets' in the explore cell while the live_train cell keeps bare 'import anywidget'/'import traitlets'; then reference 'explore_anywidget.AnyWidget' / 'explore_traitlets.Int' in the class body. The alias is a unique top-level name (no collision) and public (no underscore). Distinct from the 'mo'-alias rule (section 49: one imports cell owns it, others reference bare via reactive propagation) — that works for 'mo' because it is used at RUN time; anywidget/traitlets needed for SUBCLASSING cannot propagate that way. (2) MARKDOWN-ITALIC FALSE POSITIVE when scanning for leading-underscore identifiers — auditing a marimo notebook for accidental '_name's via regex '(?<![\w.])(_[a-zA-Z]\w*)' or grep '\b_[a-zA-Z]' over cell code produces FALSE POSITIVES inside 'mo.md(...)' strings that use '_italic_' markdown formatting (e.g. 'mo.md(f"_Loaded **{lm_name}** ... _.")' matches '_Loaded'). These are markdown EMPHASIS markers, not Python identifiers. When the scan reports a hit, verify it is NOT inside an mo.md(...) / string literal before treating it as a real identifier; cleanest fix is to drop the italic underscores from the prose (use '**bold**' or no emphasis) so the notebook is underscore-free in BOTH code AND rendered markdown.
- THIRD RESOLUTION for the unique-variable + no-underscore tension (cleanest for cells with MANY locals): wrap the cell's ENTIRE body in a function — 'def run_fix_experiment(): <body>; run_fix_experiment()' — so all internal names (model, step, optimizer, epoch, loss, etc.) become FUNCTION-LOCAL and never enter marimo's reactive graph. They cannot collide across cells (they are not reactive variables at all) and need no underscore and no per-variable prefix. The cell returns its PUBLIC output (figure, final metric) via the function's return value. Use this for HEAVY computation cells (training loops, multi-step experiments, plotting pipelines) where per-variable prefixing would mean renaming a dozen+ names by hand — one function wrapper isolates them all at once. This is also a GENERAL unique-variable mitigation (isolates locals regardless of the underscore question), complementing strategies (1) per-variable prefixing and (2) import-alias renaming for class-definition-time libs above. Discovered alphaxiv-marimo-competition-submission (07-07): the hero/heatmap/live_train/fix cells were each wrapped in a single run_*() function so internal names like model/step/optimizer/epoch could not collide with same-named locals in sibling cells.
## live-chart-vs-static-plotly
- When the user wants to visualize a training loop's loss/perplexity curve, do NOT default to an anywidget live chart. Prefer a simple STATIC plotly go.Scatter in the cell AFTER the training cell, reading the accumulated (step, loss) points stored in a public variable. The anywidget + traitlet + _esm approach documented elsewhere in this skill is the correct TECHNIQUE for genuine live/real-time updates (a marimo cell runs synchronously and cannot re-render mid-loop), but it adds significant complexity and debug surface (_esm JS, traitlets, async widget consolidation, mo.output.replace plumbing) that is NOT justified when the user just wants to SEE the resulting curve. TRIGGER PHRASES that mean 'drop the live chart, use static plotly': 'live curve doesn't have to be live anymore', 'doesn't have to be live', 'delete the anywidget one', 'just show a plot of X'. DEFAULT PATTERN: (1) training cell runs the synchronous loop, accumulates a list of [step, loss] (or [step, perplexity]) tuples into a PUBLIC non-underscore variable (e.g. live_points — see underscore-prefix-cell-local: an underscored name will NOT propagate to the plot cell); (2) a NEW cell right after it builds fig = go.Figure(go.Scatter(x=steps, y=losses)); fig.update_layout(title='Training curve'); return fig (marimo renders plotly figures natively — no matplotlib backend needed). CORRECTS the over-engineering reflex triggered by this skill's own description ('build an anywidget + traitlet whose _esm redraws...') — that is the live-update technique, presented as the answer, but live is rarely what the user wants; reach for it only when they say so. Discovered alphaxiv-marimo-competition-submission 07-07: a ternary BitNet live_train cell shipped an anywidget LiveLossChart with custom _esm; the user said 'live curve doesn't have to be live anymore' and asked for a static plotly plot in the next cell instead. Complements the plotly-over-matplotlib guidance already in this skill (prefer go.Figure/plotly.express for native interactive rendering).
## edit-cell-can-reset-name-to-underscore
- - SIBLING/REFINEMENT of edit-cell-content-leaves-stale-name (which says omitting name= leaves the name UNCHANGED): a named cell's name can SILENTLY RESET TO '_' (underscore) after a code_mode ctx.edit_cell, even when you do NOT pass name=. Observed alphaxiv-marimo-competition-submission 07-07: a cell named 'fix' became '_' after an edit_cell that rewrote the chart body (via the regex/re.sub fallback path), breaking the CellTour step that referenced it by name '{"cell_name": "fix"}'. The exact trigger is murky (likely a re.sub fallback or a prior create_cell path), but the OBSERVABLE behavior is reliable: a name can vanish to '_' across an edit. VERIFICATION + FIX: after edit_cell on any cell referenced BY NAME elsewhere (a CellTour step, a sibling cell via ctx.cells['<name>'], a disk 'def <name>'), re-read ctx.cells[...].name (or grep the disk 'def <name>') and CONFIRM the name survived; if it shows '_' or empty, restore with an explicit ctx.edit_cell('<id>', code=<unchanged or new>, name='fix'). RECOVERY CONFIRMED: passing name='fix' to a follow-up edit_cell restored the name and the tour resolved the reference. Rule: treat edit_cell as name-mutating-by-default and re-assert the intended name on any edit to a name-referenced cell. Do NOT assume the existing edit-cell-content-leaves-stale-name rule ('omitting name= is safe, name persists') holds in every code_mode path.
## latex-backslash-escape-in-md-cells
- LaTeX backslash-escape corruption in mo.md() cells: when a marimo markdown cell embeds LaTeX commands in a NON-raw Python string, Python's backslash-escapes silently corrupt the LaTeX. The recurring culprit is '\t': '\text{...}' becomes TAB+'ext{...}' and '\times' becomes TAB+'imes', so the rendered markdown shows a stray tab/garbled command where the LaTeX should be. Same class also bites '\newline' (-> newline+'ewline'), '\beta' ('\b' is BACKSPACE -> backspace+'eta'), '\sigma' (fine, no escape). SYMPTOM: a mo.md() cell with inline math renders broken LaTeX (a literal tab character or a missing leading letter where '\text{skip}' / '\times' was), with NO Python error — the string is valid, just semantically wrong. FIX: wrap any mo.md() / mo.md(f"...") cell body containing LaTeX in a RAW string (r"..." / r'''...''' / rf"...") so backslash-escape sequences pass through literally. This is distinct from code-mode-triple-quote-collision's RAW-STRING ESCAPE SUBTLETY (that note warns that raw strings PRESERVE backslash-escapes literally, which breaks escaped triple-quotes — the OPPOSITE failure; HERE the literal-preservation is exactly what you WANT for LaTeX). Rule of thumb: any marimo markdown cell containing a backslash (LaTeX math, regex shown as text, Windows paths) should use a raw string; the f-string variant rf"..." preserves both interpolation and backslashes. Discovered alphaxiv-marimo-competition-submission 07-08: eff_intro '\text{skip}' rendered as a tab, and the hook cell had a latent '\times'-becomes-tab bug; both fixed by switching to raw strings.
- rf-STRING BRACE GOTCHA (the flip side of the raw-string fix above): once you switch a LaTeX-bearing mo.md() cell to an rf-string (as recommended above to prevent backslash-escape corruption), LaTeX braces are still subject to f-string FIELD interpretation — the f-prefix means ANY '{...}' is parsed as a Python expression field, including the '{Lognormal}' inside '\mathrm{Lognormal}', the '{skip}' inside '\text{skip}', or '{a}'/'{b}' inside '\frac{a}{b}'. SYMPTOM: a misleading 'NameError: name \'Lognormal\' is not defined' (or KeyError, or a stray-brace SyntaxError) at the mo.md(rf"""...""") line, with the traceback pointing INTO the f-string — it looks like a missing variable/import, NOT an f-string brace issue, so you waste turns hunting for a typo'd name. FIX: double the LITERAL LaTeX braces so they survive f-string parsing as one literal brace: '\mathrm{{Lognormal}}', '\frac{{a}}{{b}}'. KEY DISTINCTION: only double braces that are NOT intended as f-expression fields — your intentional interpolations like '{jit_recipe_params/1e6:.1f}M' or '{n_steps}' stay single-braced. Quick audit: scan the rf-string for every '{' that is NOT followed by a Python expression you want to interpolate, and double each. This is the LaTeX-specific sibling of the f-string-brace-escaping lesson for JS/JSON object literals (memory: avoid f-strings entirely via template+sentinel+replace because nested braces are intractable); for LaTeX the brace count is usually small and tractable, so doubling is the cleaner fix. Applies beyond marimo to ANY Python f-string containing LaTeX (matplotlib '$\mathrm{...}$' titles, plotly titles, Jupyter display Math). Discovered alphaxiv-marimo-competition-submission 07-08: a mo.md(rf"""...""") cell containing '\mathrm{Lognormal}' in a sigma-distribution formula threw 'NameError: name \'Lognormal\' is not defined' — the '{Lognormal}' was parsed as an f-expression, not literal LaTeX; fixed by doubling to '\mathrm{{Lognormal}}'.
## gpu-vram-budgeting-across-cells
- Resident models/tensors from PRIOR notebook cells count against a LATER cell's VRAM budget. A big-model training/inference cell that fit fine in an ISOLATED scratch process (e.g. 3B-model STE fine-tuning peaking at ~80GB) will OOM in the notebook context because resident globals from earlier cells (lm_model, recovered_model, recovered_fp_model — often several GB each) are still held in the SAME Python process / SAME CUDA context. SYMPTOM: a cell that worked in a standalone 'uv run'/'pixi run python -c' prototype OOMs at the first training step when run in the notebook, failing by just a few MB (e.g. 'tried to allocate 44MB, 9MB free'). The margin that fit in isolation is consumed by co-resident models.
FIX: at the start of the big-model cell (or a helper function it calls), explicitly FREE the resident globals before loading the new big model: 'global recovered_model, recovered_fp_model, lm_model; del recovered_model, recovered_fp_model, lm_model; import gc; gc.collect(); torch.cuda.empty_cache()'. The 'global' + 'del' of notebook globals from inside a function works (it mutates the module namespace). After the big-model cell runs, those globals are gone — so any LATER cell that references them must reload/rebuild them, or you accept they're single-use. Alternative: reduce the big-model cell's batch/seq_len, but the weights+grad+optimizer-state memory is batch-INDEPENDENT and dominates, so freeing resident models is the higher-leverage fix.
CRITICAL DISTINCTION from scratch-cell-cuda-poisons-shared-kernel (that section is a FATAL device-side assert that corrupts the CUDA context irrecoverably and mandates a kernel RESTART): a plain CUDA OOM is RECOVERABLE — freeing memory (del + gc.collect + empty_cache) restores the ability to run GPU ops WITHOUT a kernel restart, as long as no device-side assert fired alongside the OOM. If you see a cascade where EVERY GPU cell fails after one OOM, suspect an assert poisoned the context (see transformer-context-length-cuda-assert / scratch-cell-cuda-poisons-shared-kernel), not the OOM itself. Diagnostic: 'torch.cuda.is_available()' and a tiny 'torch.zeros(1).cuda()' succeeding after the free = context healthy; that same tiny op failing = context poisoned, restart needed. So: OOM alone = free and retry; OOM-then-everything-fails = poisoned, restart.
GENERALIZES to any marimo/Jupyter GPU notebook that loads progressively bigger models across cells: budget VRAM as the SUM of all resident tensors, not just the current cell's. When escalating model size (Pythia-160M -> 410M -> 1B -> 3B), free the prior model before loading the next. Discovered alphaxiv-marimo-competition-submission 2026-07-08: a Qwen2.5-3B STE-recovery cell OOM'd at ~95GB in the notebook context (resident Pythia models + distilgpt2 approx 3.5-7GB extra) after the same recovery fit at ~80GB peak in an isolated prototype.
## vstack-return-unpack
- When a marimo cell function ends with 'return mo.vstack([...])', it returns a SINGLE _FlexContainerHtml object — NOT a tuple or list. Unpacking its return as 'a, b = func()' raises 'TypeError: cannot unpack non-iterable _FlexContainerHtml object'. This commonly bites when you refactor a cell function that ORIGINALLY returned a 2-tuple '(mo.vstack([...]), results)' into returning just the layout — but the call site still does 'output, results = func()'. The marimo error references the framework-specific _FlexContainerHtml class, which makes it look like a marimo-internal bug rather than a plain Python unpacking mismatch. Fixes: (1) if you only need the rendered layout, call the function bare ('run_fix_experiment()' — the vstack IS the cell output, no unpack needed); (2) if you need BOTH the layout AND data (e.g. results dict for a downstream cell), return an explicit 2-tuple 'return (mo.vstack([...]), results)' and keep the unpack. Diagnostic: the tell is that the function's heavy work all completed (e.g. 9 model generations ran successfully, visible in logs) and the error fires only on the FINAL unpack line — the function body is fine, only the return-shape/call-site contract disagrees. Distinct from the unique-variable-rule family (those are assignment-target collisions); this is a return-value-shape mismatch between a function and its caller within a single cell. Discovered alphaxiv-marimo-competition-submission (07-08): run_fix_experiment() returned mo.vstack([...]) but the cell did 'fix_output, fix_results = run_fix_experiment()' — TypeError on the unpack despite all 9 generations having completed successfully.
## ui-value-same-cell-cycle
- marimo FORBIDS referencing a UI element's .value in the SAME cell that defines it — it creates a CYCLE in the reactive graph (the cell would depend on its own output). SYMPTOM: a cell that does 'cfg = mo.ui.slider(1.0, 5.0, value=2.5)' and then later in the SAME body references 'cfg.value' to drive downstream computation raises a CycleError (or silently misbehaves). This is NOT a unique-variable violation — the name is defined once — it is a REACTIVE DEPENDENCY cycle: the cell that PRODUCES cfg also CONSUMES cfg.value. The instinct that bites: defining interactive sliders + reading their value + doing work, all in one cell (e.g. a generate cell that creates a CFG slider and immediately uses cfg.value to sample images). FIX: split into TWO cells — (1) a CONTROLS cell that defines+returns the UI element(s) (e.g. jit_gencfg: jit_cfg = mo.ui.slider(...); jit_nsteps_in = mo.ui.slider(...); return (jit_cfg, jit_nsteps_in)) and displays them; (2) a DOWNSTREAM CONSUMER cell that references the UI elements as reactive dependencies and reads their .value (e.g. jit_generate referencing jit_cfg.value, jit_nsteps_in.value). When the slider changes in the controls cell, marimo reactively re-runs ONLY the consumer cell — the grid/calculation regenerates. PRE-FLIGHT: when writing a marimo cell that both creates a mo.ui.* element AND uses its .value, immediately split it into controls + consumer. The controls cell must come BEFORE the consumer in cell order so the reactive edge points the right way. Discovered alphaxiv-marimo-competition-submission (07-08): a jit_generate cell defined cfg/nsteps sliders and read .value in the same body to sample a diffusion grid; split into jit_gencfg (sliders) + jit_generate (reads .value, regenerates). Sibling of the Widget-consolidation rule (that says put anywidget state+display in ONE cell for INTERNAL state management) — the two-cell split here is for marimo's OWN UI elements where .value creates a reactive cycle; they are not in tension because anywidget internal state is not a marimo reactive dependency, whereas mo.ui.*.value IS.
## competition-notebook-cache-portability
- LOCAL caches created during marimo notebook DEVELOPMENT (model checkpoints saved to /tmp or a sandbox volume, downloaded datasets, precomputed tensors) do NOT travel with the exported or shared notebook. A fresh run on a judge's / reviewer's machine re-executes EVERYTHING from scratch: re-downloads the dataset (~3 min for CIFAR), re-trains the model (~3.6 min for 8000 steps), re-runs dimension sweeps (~2-3 min). Total cold-run time can exceed 10 minutes, which judges will not wait for. The "pretrain, cache, replay" pattern (documented elsewhere in this skill) only works for the AUTHOR's local kernel — the checkpoint lives at a path (e.g. /tmp/jit_ckpts) that does not exist on the judge's machine. This is the COMPETITION/PUBLIC-SHARING sibling of library-code-changed-kernel-stale (that covers in-session import caches; THIS covers cross-machine cache absence).
- PRIMARY FIX — the HTML export IS the share artifact: run `marimo export html notebook.py -o notebook.html` (or set marimo's App auto_download to include "html"). The HTML snapshot captures the RENDERED state — figures, tables, widget outputs, markdown — exactly as they appeared on the author's machine, WITHOUT re-running any cell. Judges viewing the HTML see everything instantly. The .py still re-runs from scratch if opened in marimo, but for a competition asking for a "shareable notebook," the HTML with pre-rendered outputs is the primary deliverable. VERIFY: open the exported HTML in a fresh browser and confirm every figure/grid/widget renders with REAL content (not an empty placeholder from an un-run or stale cell).
- SECONDARY FIX — make the .py re-run tolerable: reduce training so a cold run completes in ~1-2 min (fewer steps ~2000 vs 8000, smaller model, dataset subset). Tradeoff: degrades the quality of demonstrated results, so prefer the HTML-export approach when full-quality outputs matter for judging.
- PRE-FLIGHT before declaring a competition notebook done: (1) confirm the exported HTML shows all cells rendered with real content; (2) time a cold .py Run All on a clean kernel to know what a judge opening the source will experience; (3) confirm any load-checkpoint logic ("if exists, load; else train+save") does not SILENTLY fail on a machine where the path is unwritable — it should gracefully fall back to training. Discovered alphaxiv-marimo-competition-submission 2026-07-08: the JiT notebook cached diffusion-model checkpoints to /tmp/jit_ckpts during development; the cache existed only on the author's sandbox kernel, so a judge's fresh run would re-download CIFAR + re-train the diffusion model (~10 min total) unless the HTML export carries the pre-rendered outputs.
## mo-html-does-not-render-latex-math
- mo.Html() renders RAW HTML and does NOT process markdown or $...$/$$...$$ LaTeX — only mo.md() renders math (via KaTeX). Any LaTeX written inside a mo.Html(f"""...""") styled box appears as LITERAL text in the rendered cell (e.g. '$x_0$' and '$\varepsilon$' print verbatim, including the dollar signs), with NO error. SYMPTOM: a custom-styled box (gradient bg, pill badges, bordered callout) built with mo.Html shows un-rendered math like '$x_0$' / '$\sigma$' where you expected subscripts/Greek letters — the prose reads correctly but the math is raw. FIX (two options): (1) CONVERT the math to HTML entities/tags within the mo.Html string — '$x_0$' -> 'x<sub>0</sub>', '$x_t$' -> 'x<sub>t</sub>', '$\varepsilon$' -> 'ε', '$\sigma$' -> 'σ', '$\times$' -> '×', '$\approx$' -> '≈', '$\hat{x}$' -> '<i>x̂</i>' (combining circumflex) or 'x̂'; display math $$...$$ -> a centered styled <div> with the same entity-encoded content. (2) RESTRUCTURE so the math-bearing prose lives in a mo.md() cell (which DOES render LaTeX) and the styled chrome is pure layout — e.g. mo.vstack([mo.md(r'...prose with $x_0$...'), ...], ...) wrapped in a container, or mo.md(...).callout_style. Decision rule: when a box needs BOTH custom CSS AND inline math, either inline-convert the math to HTML entities (small math count, quick) or move the prose to mo.md and keep the box as layout (large math count). DISTINCT from the two sibling sections already in this skill: latex-backslash-escape-in-md-cells (mo.md backslash corruption, raw-string fix) and the rf-string-brace gotcha (mo.md(rf'...') brace doubling) — THOSE are about mo.md() rendering math CORRECTLY with escaping care; THIS is about mo.Html() not rendering math AT ALL. When authoring a styled box, decide upfront whether it is mo.Html (math->entities) or mo.md (math stays as LaTeX). Discovered alphaxiv-marimo-competition-submission 07-08: close.py hero/arc/limits box (mo.Html f-string) had '$x_0$' / '$\varepsilon$' rendering as literal text; fixed by converting to 'x<sub>0</sub>' and 'ε'. Same notebook's hero.py, cifar_data.py, manifold_game.py all use mo.Html and would silently mangle any LaTeX left inside them.
## mo-callout-needs-mo-md-wrap-to-render-markdown
- mo.callout() does NOT render markdown on a PLAIN STRING — it wraps the string in a bare `<span>` and passes it through as LITERAL text, so `**bold**`, `*italic*`, `$...$` LaTeX, AND HTML named entities (`—`, `→`, `×`) ALL appear verbatim with NO error. Verified empirically (marimo 0.23.13): `mo.callout('**bold** — $x_0$')` -> `<span>**bold** &mdash; $x_0$</span>` (literal asterisks, double-escaped entity, raw dollar-math). ROOT CAUSE: `mo.callout(value)` internally calls `as_html(value).text`; for a plain string `as_html` wraps it in a `<span>` with NO markdown processing — only `mo.md()` invokes the markdown+KaTeX renderer. DEFINITIVE FIX: wrap the callout content in `mo.md(...)` — `mo.callout(mo.md('**bold** — $x_0$'), kind='success')` renders correctly: `**bold**` -> `<strong>bold</strong>`, `$x_0$` -> `<marimo-tex>...x_0...</marimo-tex>` (KaTeX), and `—` -> `—` (correct entity, displays as —). This single wrap fixes BOTH the markdown AND the LaTeX in one shot. HTML NAMED ENTITIES still need Unicode substitution even inside mo.md (the markdown renderer passes `—`/`→`/`×`/`ε` through literally) — replace with actual Unicode chars: `—`->'—', `→`->'→', `“`->'\u201c', `”`->'\u201d', `×`->'×', `ε`->'ε', `σ`->'σ' (Python source is UTF-8, accepts these directly). CORRECTED RENDERING-CAPABILITY MATRIX: mo.md(string) = markdown yes, LaTeX yes, HTML named entities NO (use Unicode). mo.callout(plain string) = markdown NO, LaTeX NO, entities NO — NOTHING renders (wrap in mo.md to get mo.md's behavior). mo.callout(mo.md(string)) = markdown yes, LaTeX yes, entities NO (inherits mo.md). mo.Html(string) = raw HTML yes (tags + entities), markdown NO, LaTeX NO. SYMPTOM: a callout shows visible `**asterisks**`, visible `$x_0$`/`$\varepsilon$`, AND visible `—`/`→` all at once — the telltale sign that the content was passed as a plain string, not wrapped in mo.md. When ALL THREE (bold + LaTeX + entities) fail in a callout, the fix is `mo.callout(mo.md(...))`, NOT per-symbol Unicode/LaTeX substitution. PREVIOUS (07-09) diagnosis of this same symptom was INCOMPLETE: it concluded 'entities don't render + LaTeX needs restructure' and applied symptom-level Unicode/HTML-entity substitution, MISSING that markdown bold ALSO failed and that the single root cause is the missing mo.md() wrap — re-encountered 2026-07-09 on manifold_ext/denoise_ext/jit_generate/jit_train/denoise_gallery/dim_sweep callouts (6 callouts all passing plain f-strings, none wrapped in mo.md). IMPORTANT INTERACTION with the mo.Html section above: the 'convert LaTeX to HTML entities' workaround (`ε`, `σ`) works ONLY in mo.Html() (raw HTML context where the browser processes entities) — it does NOT work in mo.md()/mo.callout() (markdown context). Sibling sections: latex-backslash-escape-in-md-cells (mo.md backslash corruption, raw-string fix), rf-string-brace gotcha (mo.md(rf'...') brace doubling), mo-html-does-not-render-latex-math (mo.Html math->entities). Distinct from all three: THIS is about mo.callout needing a mo.md() wrap to render ANY markdown at all.
## app-config-not-settable-via-code-mode
- marimo's APP-LEVEL CONFIG (width, auto_download, etc. — the marimo.App(width='medium', auto_download=['html']) in the notebook file header) has NO code_mode API. ctx (AsyncCodeModeContext) exposes cell-level operations (edit_cell, create_cell, move_cell, run_cell) but NO method to set app-level config like width or auto_download. Probed via dir(ctx) — no app-config/config/set_width/set_auto_download method exists. CONSEQUENCE: when a rubric or requirement asks for width='medium' and auto_download=['html'/'python'] (common for competition submissions — see competition-notebook-cache-portability which notes auto_download carries the pre-rendered HTML), you CANNOT set it programmatically via code_mode. The app config lives ONLY in the notebook file header (marimo.App(...)), which code_mode does not mutate. OPTIONS: (1) ASK THE USER to set it in the marimo UI (gear/settings menu → App config → set width and auto_download) — this is the only reliable path during a live code_mode session because marimo re-serializes the file on save and would clobber a manual file edit (see main-block-clobber-on-save); (2) if the session is NOT running (STOP-REWRITE-RESTART path), edit the file header directly — app = marimo.App(width='medium', auto_download=['html']) — then start the server. Do NOT waste turns probing dir(ctx), dir(cm), or searching for a programmatic setter — verified absent in marimo >=0.13. Discovered alphaxiv-marimo-competition-submission 2026-07-08: the JiT competition notebook needed width='medium' + auto_download=['html'] for a rubric item; multiple turns spent checking ctx methods before concluding the UI is the only path during a live session.
## env-inheritance-after-env-rewire
- Sibling of library-code-changed-kernel-stale (that's the kernel's PYTHON IMPORT cache going stale after you edit a .py the notebook imports). THIS section is the ENV-VAR analogue: after you rewire .env to a new endpoint (switching vLLM serving endpoint ericmjl -> nll-ai workspace, or any LLM_MODEL / TUTORIAL_LLM_BASE_URL / *_API_KEY change), a RUNNING marimo server keeps the OLD endpoint values — because python-dotenv's load_dotenv() does NOT override existing env vars by default (override=False), and the server inherited the stale SHELL env at launch, so load_dotenv silently no-ops on keys already in os.environ. .env on disk is NOT ground truth for the running process. SYMPTOM: cell output or resolved config still prints the old MODEL/BASE even though 'cat .env' shows the new values. FIX: restart the marimo server in a CLEAN shell — unset the stale vars in the SAME command before relaunching so the child does NOT inherit them: 'kill <pid>; sleep 1; unset LLM_MODEL TUTORIAL_LLM_BASE_URL TUTORIAL_LLM_API_KEY OPENAI_API_KEY; cd <repo> && nohup pixi run marimo edit notebooks/<f>.py --port <p> --no-token > /tmp/marimo.log 2>&1 & echo relaunched pid $!'. Verify via the restarted process's OWN resolved config (run a cell printing MODEL/BASE or print(os.getenv(...))), NOT by re-reading .env on disk. The unset-in-same-command trick matters because a persistent bash tool shell keeps the unset for future calls too, but chaining kill+unset+launch in ONE call guarantees the child inherits the cleaned env. GENERALIZES to any running Python process using load_dotenv (jupyter, a long-lived FastAPI/CLI service), not just marimo. Distinct from library-code-changed-kernel-stale (import cache — fixable via importlib.reload or next natural restart) and from STOP-REWRITE-RESTART (that's a deliberate notebook RESTRUCTURE on disk): THIS is purely about the server's INHERITED ENV, fixable by a clean-env restart with no notebook file change. Discovered build-deep-research-agent 07-08 switching the notebook from ericmjl to nll-ai vLLM endpoint.
- NO-RESTART COMPLEMENT (discovered build-deep-research-agent 07-09): the entry above covers vars ALREADY in os.environ (shell-inherited at launch) — load_dotenv() no-ops, restart with unset. But when the user ADDS NEW vars to .env that are ABSENT from os.environ (e.g. user just pasted ZOTERO_LIBRARY_ID/KEY mid-session, then says "I set env vars in .env, try again"), load_dotenv() WILL add them even with override=False (absent keys are always populated), so NO restart is needed. Run `from dotenv import load_dotenv; load_dotenv(override=True)` in a test cell to pick up the new creds in the running kernel — use override=True defensively in case any var is partially set. This is the faster path for the common "I added my API key to .env, try again" workflow; reserve the unset+restart path for vars that were already in the shell env. Decision rule: if the var was absent at kernel launch (not shell-inherited), no-restart works; if it was shell-inherited and changed, you must unset+restart.
## hide_code-rubric (competition/pedagogical notebooks)
- COMPETITION + PEDAGOGICAL NOTEBOOK RUBRIC: ~75% of cells should be hide_code=True, keeping ~25% (the PEDAGOGICAL 'real code' cells) VISIBLE so judges/readers see the actual method. Do NOT set hide_code=True on 100% of cells — a fully-hidden notebook scores LOWER on the 'code is visible / pedagogical value' axis even when every cell has a markdown narrative. A 100%-hidden notebook is a common end-state after batch-hiding every cell for a clean look, but it under-performs the rubric. Target ~3-4 VISIBLE cells out of ~15 (~75-80% hidden).
- DECISION HEURISTIC for WHICH cells to keep VISIBLE (flip hide_code=False), in priority order: (1) the METHOD / model-definition cell (e.g. jit_recipe = the DiT+EDM architecture — the paper's actual method, the single most important cell to show); (2) the THESIS-EXPERIMENT cell (e.g. dim_sweep = the x0-vs-eps patch-size sweep that IS the notebook's central empirical claim); (3) the TRAINING loop (e.g. jit_train). Keep HIDDEN: hero/banner SVG art, CellTour/tour widgets, pure-import + DEVICE-setup cells, UI/slider/config cells, gallery cells. Discovered alphaxiv-marimo-competition-submission (07-08): all 15 cells were hide_code=True (100%), flagged in self-review; flipped hide_code=False on jit_recipe + dim_sweep + jit_train via ctx.edit_cell(name, code=None, hide_code=False).
- TECHNIQUE: ctx.edit_cell(name, code=None, hide_code=False) flips the flag WITHOUT disturbing the cell body (see hide_code-verification above). Verify against the ON-DISK file (@app.cell(hide_code=False)) or the edit_cell 'edited code and config' confirmation — NOT the unreliable Cell-view attribute probe.
- SELF-CONTRADICTION CHECK (high-value, easy to miss): if a cell's own PROSE claims the code is visible ('The cells below keep this code visible because it is the paper's actual method') but the cell is marked hide_code=True, that is a direct self-contradiction — flip it. When reviewing a notebook, grep each cell's narrative for 'visible' / 'shown below' / 'we keep' / 'the code below' and reconcile every such claim against the actual hide_code flag.
## skip-validation-trap
- DISTINCT FROM skip_staleness_check=True (the safe staleness bypass): when a create_cell/edit_cell batch is rolled back by a MultipleDefinitionError (a name defined in two cells), the traceback's suggestion — 'To skip validation, use: async with cm.get_context(skip_validation=True) as ctx' — names a DIFFERENT, TRAP parameter. skip_validation=True suppresses the DRY-RUN COMPILE that detects name collisions before flush, so the batch APPEARS to land, BUT marimo STILL enforces unique cell-scoped names at RUNTIME — the second defining cell is marked 'marimo-error' (output suppressed) and downstream consumers hit NameError when the notebook actually runs. So suppressing validation produces a silently-broken notebook, not a working one. ALWAYS fix the colliding names instead (namespace by concern — docstore -> zotero_docstore, payload -> zotero_semantic_payload, etc.), never pass skip_validation=True to dodge a collision. Two separate skip flags, two separate purposes: skip_staleness_check=True = safe mitigation for StaleCellError on edit_cell (batch still flushes cleanly); skip_validation=True = trap that masks a real defect (batch lands but notebook is broken). Decision rule: if the dry-run failure is a MultipleDefinitionError, rename the variable; if it is a StaleCellError, pass skip_staleness_check=True. Discovered build-deep-research-agent Part 3 (07-09): ex5 cells reused docstore/side_table/payload/zotero_payload owned by other cells; the skip_validation suggestion in the traceback was correctly rejected in favor of namespacing the ex5 vars (zotero_docstore, zotero_side_table, zotero_semantic_payload).
## run-cell-blocks-kernel-synchronous
- ctx.run_cell executes SYNCHRONOUSLY on the marimo kernel — a cell that takes minutes (heavy docstore builds, PDF downloads for 90+ items, mass embeddings of 6000+ chunks) blocks ALL subsequent execute-code.sh calls for its full duration. The execute-code.sh timeout is CLIENT-SIDE (your HTTP wait expired); the kernel keeps running the cell regardless. If the cell hangs (vs just slow), the kernel becomes permanently WEDGED and needs a restart via the marimo UI (no reliable code_mode kernel-restart API exists — ctx.execute_command has no documented 'restart kernel' command). PRACTICAL RULE: for heavy operations taking minutes, test the path via direct Python OUTSIDE the notebook first (pixi run python -c '...' or a standalone uv run script.py calling part3.build_zotero_docstore()), NOT via ctx.run_cell. A direct reference call confirms the code path works just as well as a cell re-run, without blocking the kernel. Only use ctx.run_cell to confirm the CELL ITSELF works if you specifically need to validate the cell wiring, and accept the kernel will be unavailable for the full duration — you CANNOT poll or status-check meanwhile (your execute-code calls queue behind the running cell and time out). WORKFLOW: (1) test heavy operations via direct Python; (2) if the direct test passes, the notebook cell (which delegates to the same function) will pass too — skip the cell re-run; (3) if you must re-run the cell, do it as the LAST operation and warn the user the kernel will be busy for N minutes. UNCONFIRMED HYPOTHESIS: re-running a cell that calls docstore.reset() on an EXISTING LanceDB table (6000+ chunks from a prior direct test) then re-appends may be significantly slower or hang, whereas a fresh build completes in ~200s — the reset/re-append path on existing data was not confirmed as the cause but is the leading hypothesis for why the cell re-run wedged at 800s+ when the direct fresh build took 204.7s. Distinct from code-mode-edit-flush-on-exit (stale reads within one block) and mid-batch-exception-discards-flush (exception aborts the flush) — THIS is about the EXECUTION COST of the cell itself blocking the single-threaded kernel, not about the flush mechanics. Discovered build-deep-research-agent Part 3 (07-09): direct part3.build_zotero_docstore() call completed in 204.7s (94 items, 6491 chunks, semantic search returned relevant results); re-running the ex5_build cell via ctx.run_cell wedged the kernel for 800s+ — four subsequent execute-code calls (120s, 300s timeouts) all timed out waiting for the kernel; the direct test had ALREADY confirmed the path, making the cell re-run unnecessary.
## UI element value setting via code_mode (set_ui_value gotcha)
- To programmatically set a mo.ui.* element's value from code_mode (e.g. to 'click' a mo.ui.run_button gated behind mo.stop(not btn.value, ...)), use ctx.set_ui_value(element, value). CRITICAL: set_ui_value takes the actual UI element OBJECT (it reads element._id internally), NOT a string name. Fetch the object from ctx.globals first: 'btn = ctx.globals["run_ex1"]; ctx.set_ui_value(btn, True)'. Passing ctx.set_ui_value("run_ex1", True) silently does nothing. After setting the value, call ctx.run_cell(name) to re-execute the gated cell. NOTE: exercise cells using mo.stop(not run_exN.value, ...) display status 'exception' with a placeholder message until the button is activated — this is an EXPECTED halt (mo.stop), not a real error. The repo-local marimo-pair SKILL.md mentions set_ui_value(element, new_value) but omits the ctx.globals lookup and the object-not-string gotcha. Verified against marimo _code_mode/_context.py:1362 (set_ui_value does UIElementId(element._id)). Discovered build-deep-research-agent exercise cells 07-10.
## set_ui_value triggers reactive re-runs
- GOTCHA: `ctx.set_ui_value(element, value)` inside code_mode DOES trigger downstream reactive cell re-execution — unlike setting an anywidget trait directly (`slider.value = 5`) which does NOT fire the reactive graph. If downstream cells consume the UI element's value and do expensive work (LLM calls, model inference, large computations), the code_mode block will TIME OUT waiting for them (symptom: values print as set successfully, then the block hangs). FIX: after calling set_ui_value on elements with expensive downstreams, exit the code_mode block promptly and poll cell statuses in a subsequent execute-code call rather than waiting inside the same block. ACCESS PATTERN: the `element` arg comes from `ctx.globals['element_name']` — ctx.globals is a dict of all top-level kernel names including mo.ui.* element objects. Discovered build-deep-research-agent 2026-07-10 programmatically activating 3 mo.ui.run buttons whose downstream cells called the LLM.
## generated-with-version-downgrade-on-save
- When a marimo notebook file was originally generated by a NEWER marimo version (e.g. __generated_with = '0.23.13') and you open+edit it with an OLDER running marimo version (e.g. 0.23.8), the kernel's auto-save RE-STAMPS __generated_with to match the RUNNING version (0.23.8), creating a version DOWNGRADE string in git diffs. This is harmless metadata noise — the field tracks which marimo version last serialized the file, nothing more — but it appears alongside your meaningful diff and can be confusing if you don't recognize it. The same save event may also strip trailing whitespace (ruff format on save) and normalize other formatting. When committing, recognize the __generated_with version-string change as a re-serialization artifact, NOT an intentional edit; do not waste reasoning turns deciding whether to split it into a separate commit. Sibling of main-block-clobber-on-save (same re-serialization mechanism, different artifact: version header vs __main__ block). Discovered build-deep-research-agent notebooks 07-10: editing a PR's notebook (generated with 0.23.13) under local marimo 0.23.8 downgraded the version string on kernel save.
## ui-element-value-setting-via-code-mode
- DIAGNOSTIC NUANCE: mo.stop() called WITHOUT a second message argument (mo.stop(condition) not mo.stop(condition, mo.md('...'))) produces EMPTY cell.output.data — there is no placeholder message to hint that mo.stop is the cause. This makes the 'status=exception' look MORE like a real error than the with-message case (which shows a visible placeholder). When diagnosing a cell with status=exception and empty output.data: (1) read the cell code for mo.stop() — if present and the condition is True at the current reactive state, it is an expected halt; (2) check the condition's dependencies — e.g. mo.stop(not save_env.value) stops when the button hasn't been clicked, mo.stop(not run_exN.value) stops until the exercise run button is activated; (3) the FORM/UI cell defining the UI elements should be status=idle — if it is, the downstream validation cell's exception is expected, not broken. Discovered build-deep-research-agent startup_validation cell 07-10: mo.stop(not save_env.value) with empty output.data initially looked like a real exception before tracing the mo.stop call.
## conditional-rendering-if-block
- In marimo, ONLY the LAST TOP-LEVEL expression of a cell renders as output. An expression placed inside an 'if' block (or for/with/try-except) is a nested STATEMENT — it evaluates and its result is DISCARDED, NOT captured as the cell's output. This is the FUNDAMENTAL marimo rendering model: marimo inspects the cell function body's AST and captures only the last EXPRESSION STATEMENT at the MODULE/TOP level of the cell (outside any control-flow block).
SYMPTOM: a cell containing:
if not env_configured:
mo.vstack([...])
shows 'no output' / 'No output data' even though the code ran, the branch was entered, and the expression evaluated — the mo.vstack result was discarded because it's inside the if, not at the cell's top level.
FIX (two patterns for conditional rendering):
(1) ASSIGN-THEN-RETURN: bind the result to a variable inside the branch, then place the variable as the cell's last top-level expression:
if not env_configured:
result = mo.vstack([...])
else:
result = None
result
In .py format, this is equivalent to making 'result' the value before 'return (result,)'.
(2) EARLY-EXIT via mo.stop: use mo.stop to halt the cell when the condition IS met, leaving the render expression as a bare top-level statement:
mo.stop(env_configured)
mo.vstack([...])
mo.stop(condition) halts the cell (no further execution) when condition is True; when env_configured is False (not stopped), mo.vstack([...]) runs as the last top-level expression and renders. This is the RENDERING counterpart of the mo.stop early-exit pattern documented in the return-outside-function section (#71). See also the mo.ui.run_button activation pattern (ctx.set_ui_value) for programmatic activation of mo.stop-gated cells.
DISTINCT FROM expression-in-try-block: bare expressions inside try/except blocks ARE captured by marimo for display (and must be suppressed with '_ = expr'). This is the OPPOSITE case: bare expressions inside if/for/with blocks are NOT captured. The asymmetry: marimo's display-capture scans try/except bodies for expressions to display, but does NOT scan if/for/with bodies — so the same mo.vstack([...]) that renders as a bare top-level expression is silently discarded when nested inside an if.
TOKEN-EFFICIENCY NOTE: when a marimo cell shows 'no output' and the cell body has the render expression inside an if/for/with block, do NOT spend reasoning turns hypothesizing about 'whether expressions inside if-blocks render' or 'kernel state issues' — the answer is definitive: they do NOT. Restructure to assign-then-return or mo.stop. Discovered build-deep-research-agent notebook startup_form cell (07-10): mo.vstack inside 'if not env_configured:' produced 'no output'; the assistant burned multiple reasoning turns before recognizing the top-level-only rendering rule.
- CORRECTION to the 'DISTINCT FROM expression-in-try-block' note below: the earlier claim that 'bare expressions inside try/except blocks ARE captured by marimo for display' has been SUPERSEDED (07-10). Memory #604 + two independent 07-10 observations (startup_form cell, Part 4 ex1_run cell) show try/except expressions are ALSO discarded — same as if/for/with. The reliable rule is UNIFORM: marimo renders ONLY the last TOP-LEVEL expression (or return value); expressions inside ANY nested block are discarded. The old try/except-vs-if/for/with asymmetry claim is no longer trustworthy.
## cross-cell-symbol-removal
- CROSS-CELL SYMBOL REMOVAL (generalization of the matplotlib-removal cascade above): when you remove a symbol — a variable, UI widget, or import — from its DEFINING cell, EVERY downstream cell that references that symbol breaks with NameError on the next reactive re-run. The safe removal procedure is NOT 'delete from the defining cell and see what breaks'; it is: (1) TRACE all consumers FIRST — grep the symbol name across all notebook cells (for code_mode: iterate `ctx.cells` and search each cell's `.code`; for disk editing: `rg` the .py file) to enumerate every cell that references the removed name; (2) EDIT ALL consumer cells in ONE coherent batch — replace each removed reference with a safe fallback (e.g. `os.getenv("REMOVED_VAR", "")` for an env var, a default value, or delete the reference entirely if the logic no longer needs it); (3) ONLY THEN remove the symbol from the defining cell. If you edit the defining cell first and then try to fix consumers one-by-one, the intermediate state has broken cells that marimo cascades errors into. This is the INVERSE of ADDING names (where the unique-variable-rule prevents collisions); removal requires the same multi-cell awareness but in reverse — proactive consumer-fixing before the defining cell is changed. Discovered build-deep-research-agent (07-10): removing `api_key` + `api_key_input` from the `startup_form` cell required also editing `startup_validation` (which used `api_key_input.value` to write `.env` and `api_key` for the env-read else-branch); the agent traced both references, replaced the removed `_key = api_key` with `_key = os.getenv("TUTORIAL_LLM_API_KEY", "").strip()` to preserve BYO-auth via .env, and edited both cells in one batch — the edit landed clean with no NameError.
- POST-EDIT 'ran (error)' TRIAGE: when marimo reports a cell status as `exception` or `ran (error)` after your edit, do NOT assume the edit broke it — READ THE ACTUAL TRACEBACK from `cell.output.data` to distinguish a STRUCTURAL error (NameError on a removed/renamed symbol = your edit broke the reactive graph) from an ENVIRONMENTAL exception (TimeoutError from an endpoint ping, ConnectionError, FileNotFoundError from a missing .env, etc. = pre-existing/environmental, not caused by your edit). The two require opposite responses: a structural NameError means you missed a consumer cell (go back to the cross-cell-symbol-removal procedure above); an environmental exception means your edit is fine and the error is orthogonal (note it to the user as pre-existing, do not try to 'fix' it by reverting the edit). Key diagnostic: if the cell ran PAST the lines you edited (the traceback originates at a network/filesystem call, not at a variable reference), the edit is structurally correct. Discovered build-deep-research-agent (07-10): after removing `api_key_input` from `startup_form` + editing `startup_validation`, the validation cell showed status `exception` — but the traceback was `TimeoutError: The read operation timed out` from a `urlopen(...)` endpoint health-check, NOT a NameError on the removed widget. The edit was structurally correct; the timeout was because the notebook was running in a worktree whose `.env` still pointed at a cold/stale endpoint. This is the sibling of DUPLICATE-DEFINITION DIAGNOSTIC (read cell.output.data for the cause BEFORE assuming a graph bug) — the same triage discipline applies to post-edit errors: read the traceback first, theorize second.
## description-trigger
- Add to the description Covers list after 'the matplotlib fig, ax plt.subplots collision': 'cross-cell SYMBOL REMOVAL (trace all downstream consumers before removing a variable/widget/import from its defining cell, edit them in one batch — the inverse of unique-variable collision rules), post-edit ran-error traceback triage (read cell.output.data to distinguish structural NameError from environmental TimeoutError/connection error after an edit)'.
- ADD to the description Covers list: 'CellOutput .errors attribute (check cell.output.errors — NOT .data — for the traceback when a cell has status=exception and .data is empty; the CellOutput API surface is: data/errors/stderr/stdout/mimetype/channel/empty/asdict/timestamp)'.
## stale-env-vars-load-dotenv-override-false
- load_dotenv(override=False) in a LONG-RUNNING kernel (marimo/Jupyter) produces STALE os.getenv() reads after the .env file is updated. Root cause: when the kernel process STARTED, its process environment was populated with the .env values that existed at startup time (or shell env vars). load_dotenv(override=False) only sets env vars that are NOT already present — so once a var is baked into the process env at startup, subsequent .env file changes NEVER reach os.getenv (override=False refuses to overwrite). Symptom signature (build-deep-research-agent notebook 01 startup_form, 07-10): the startup_form widget shows OLD vLLM values (LLM_MODEL=openai/google/gemma-4-12B-it, TUTORIAL_LLM_BASE_URL=...vllm...) even though the .env on disk has the CORRECT Ollama values (openai/gemma4:12b, ...ollama-service...) and the README defaults are also Ollama. The .env on disk is right; os.getenv returns wrong values because the process env was set at startup with the old values and override=False won't refresh them. Diagnostic procedure: (1) compare the .env file on disk (Read tool) against os.getenv() output in the kernel — a mismatch with override=False IS this bug; (2) confirm the kernel's process env by printing os.environ.get(var) directly (NOT via the cell's env_values dict, which can itself be stale due to code-mode globals-proxy timing — see code-mode-edit-flush-on-exit). Fix: change load_dotenv(env_path, override=False) to load_dotenv(env_path, override=True) in the startup cell and re-run it (ctx.run_cell) — this overwrites the stale process env with the current .env values. Alternative: restart the marimo kernel (heavier — loses live graph state). Generalizes to ANY long-running Python process (notebook kernels, dev servers, daemons) that calls load_dotenv with override=False and whose .env changes at runtime. Prefer override=True in long-lived processes where the .env is expected to change (notebooks under active development, multi-worktree setups where a worktree's .env differs from the one loaded at server start). NOTE: the kernel's CWD also matters — a multi-worktree marimo server started from worktree A will resolve env_path relative to A's CWD even when editing worktree B; verify the CWD (os.getcwd()) matches the worktree whose .env you expect. Distinct from code-mode-edit-flush-on-exit (stale READ-BACK within one code_mode block) and ctx.cells-stale-after-mutation — those are about marimo's notebook graph state; THIS is about the OS process environment being immutable under override=False.
## mkdocs-admonition-syntax-does-not-render-in-marimo
- MkDocs/Material admonition syntax — the collapsible '??? type "title"' and non-collapsible '!!! type "title"' fenced blocks followed by indented content — does NOT render in marimo markdown cells. Marimo's markdown renderer does NOT support the Python-Markdown admonitions extension, so the raw '??? question "Check your understanding"' text appears LITERALLY in the rendered output with NO error and NO styling. This is a natural carryover mistake in projects that have BOTH MkDocs docs (where admonitions work) AND marimo notebooks (where they don't) — the build-deep-research-agent tutorial repo has both, and admonition syntax written for docs/ silently fails when pasted into a marimo mo.md() cell. FIX: use mo.callout(mo.md('...'), kind='info'|'warn'|'danger'|'success') — marimo's native callout mechanism (see mo-callout-needs-mo-md-wrap-to-render-markdown for the mo.md() wrap requirement). KIND MAPPING: MkDocs 'question'/'info'/'note' -> kind='info'; 'warning'/'abstract' -> kind='warn'; 'danger'/'failure' -> kind='danger'; 'tip'/'success'/'check' -> kind='success'. When the admonition sits INSIDE a prose mo.md(dedent('''...''')) cell, you must RESTRUCTURE the cell's last expression to pull the admonition out into a separate mo.callout and combine both via mo.vstack: 'return mo.vstack([mo.md(dedent('''...main prose...''')), mo.callout(mo.md('...question + answer...'), kind='info')])'. You CANNOT call mo.callout() inside an mo.md() string — it is a Python function, not markdown syntax. Discovered build-deep-research-agent 2026-07-10: three cells (text_in_text_out in notebook 01, why_these_sources_bridge + vector_search_insight in notebook 02) had '??? question' check-for-understanding admonitions that rendered as literal text; user flagged it and directed to mo.callout().
## new-notebook-file-scaffold
- ## new-notebook-file-scaffold
- COMPLETE ON-DISK .py TEMPLATE for authoring a NEW marimo notebook from scratch (the assembly of structural-block-types + import-cell-return + notebook-file-editing escape-hatch #1 into one concrete file a future agent writes directly, discovered build-deep-research-agent 07-10). Use this when scaffolding a new notebook, extracting/splitting cells from one notebook into another, or recovering from mass-unparsable cells. Reverse-engineering this format from an existing file costs 5+ turns; this template is the shortcut. The minimal valid file structure:
```python
# /// script
# requires-python = ">=3.12"
# dependencies = [
# "marimo",
# "<other-deps>",
# ]
# ///
import marimo
__generated_with = "0.23.8"
app = marimo.App(width="medium")
with app.setup(hide_code=True):
# Shared imports/constants/functions — GLOBALLY visible to ALL cells
# WITHOUT being passed as cell params or returned (see contract below).
import marimo as mo
from textwrap import dedent
# ... other shared imports, helper functions, constants ...
@app.cell(hide_code=True)
def defining_cell():
# Defines variables other cells need → MUST return them as a tuple.
import os # cell-LOCAL deps imported here, not in setup
x = mo.ui.slider(0, 10)
result = compute(x.value)
return result, x # <-- exported symbols
@app.cell
def consuming_cell(result, x):
# Consumes variables → receives them as FUNCTION PARAMETERS
# (same names the defining cell returned).
mo.md(f"Result: {result}")
return
@app.cell(hide_code=True)
def standalone_cell():
# Neither imports from nor exports to other cells → bare 'return'.
mo.md("# Title")
return
if __name__ == "__main__":
app.run()
```
- UNIFIED CELL-DATAFLOW CONTRACT (three cases, verified against build-deep-research-agent notebook 01, 07-10). A symbol reaches a cell by exactly ONE of these three paths:
(1) app.setup GLOBAL: any name defined/imported in 'with app.setup(...):' is available to EVERY cell as a bare global — it is NOT a cell param and NOT in any return tuple. This is how 'mo', 'dedent', and shared helper functions flow. (2) DEFINING cell → return tuple: a cell that creates a variable another cell needs MUST end with 'return (var1, var2, ...)' listing ALL exported names. (3) CONSUMING cell → params: a cell that uses a name from another cell receives it as a function parameter with the SAME name ('def consumer(var1, var2):'). The defining cell's return tuple and the consuming cell's parameter list must agree on names. A cell that neither imports-setup-symbols-beyond-the-globals nor needs to export just ends with 'return'. GOTCHA: 'mo' is in setup (case 1), so cells use 'mo' with NO 'mo' param and NO 'mo' in their return tuple — this looks wrong if you expect mo-as-param, but it is correct because mo is a setup global. The 'cell-parameter-verification' rule ('mo must be received as parameter') applies ONLY when mo is NOT in the setup block (e.g. a plain def _(mo): cell); when mo IS in setup, it is global and needs no param.
- EXTRACTION / SPLIT PROCEDURE (copy cells from notebook A into a new notebook B): (1) read notebook A's disk file; (2) identify the cells to extract by their 'def <name>' and their full bodies down to the 'return' statement; (3) write notebook B with the template above: its OWN PEP 723 header (with B's deps), its OWN app.setup (at minimum 'import marimo as mo' — copy any setup symbols the extracted cells use), then paste the extracted cells VERBATIM (preserving their internal cell-local imports and their return tuples / param signatures); (4) 'if __name__ == "__main__": app.run()'. Verify: compile-check ('python -m py_compile B.py'), then confirm @spec anchors and key logic survived via grep. Open in marimo to confirm it is a valid runnable notebook. The extracted cells' param/return wiring is self-contained — if cell X returns (a, b) and cell Y takes (a, b), pasting both preserves the contract as long as B's setup provides the globals they reference.
## three-layer-cell-state-divergence
- ## three-layer-cell-state-divergence
- Marimo cell state exists in THREE layers that can all DIFFER, and confusing them is the single biggest time-sink when sourcing cell content to MOVE between branches or notebooks. The three layers: (1) LIVE KERNEL (ctx.cells — volatile in-memory; cells created/edited via code_mode that are never explicitly saved-to-disk-AND-committed vanish on restart or when the kernel re-serializes the disk file from a DIFFERENT state); (2) DISK (the .py working-tree file — what git status/diff tracks, may have uncommitted changes); (3) GIT COMMIT (git show <branch>:<file> — the durable committed state). The kernel is the WORST source of truth for durable content — it is ephemeral and can hold cells that were NEVER persisted to either disk or git.
- SOURCE-OF-TRUTH VERIFICATION for cross-branch cell extraction (discovered build-deep-research-agent 07-10): when you plan to EXTRACT cell code from branch A to create a new file/notebook on branch B, do NOT trust what you remember from the running kernel. A split-cell edit (e.g. startup_form + startup_validation created via code_mode) can exist ONLY in the kernel — never saved to disk, never committed — while 'git show branchA:notebooks/X.py' shows the ORIGINAL combined cell (or different cells entirely), and the disk file shows yet another version. DIAGNOSTIC: 'git show <branch>:<file> | rg "def <cell_name>"' reveals the COMMITTED truth; 'rg "def <cell_name>" <file>' reveals the DISK truth; ctx.cells reveals the KERNEL truth. If the kernel has a cell that NEITHER git show NOR disk has, it was never persisted — treat it as ephemeral and source from disk/commit instead. RESOLUTION when the three diverge: use the DISK or COMMITTED version that the target branch actually has (whichever matches the branch you are moving TO), NOT the volatile kernel version you edited. Discovered build-deep-research-agent 07-10: spent 5+ reasoning turns reconciling kernel (had split startup_form+startup_validation) vs disk (combined startup_validation) vs committed branch (had neither — just hero) before landing on 'use the combined disk version as the source for the new 00_check.py'.
- DISTINCT from the related hazards already documented above: (a) GIT-OPERATION CLOBBER (that is: a git merge/rebase changed the disk file UNDERNEATH a running kernel, and the kernel clobbered the git changes on its next auto-save — the fix is STOP the server before git ops); (b) DISK-FILE LAG (that is: the disk .py lags behind the kernel after a successful kernel flush, breaking pytest — the fix is stop-server-to-trigger-save); (c) PORT→WORKTREE→BRANCH VERIFICATION (that is: editing the wrong port edits the wrong branch — the fix is cross-reference port→worktree before editing). THIS hazard (three-layer divergence on cross-branch SOURCING) is different: the kernel held cells that were simply NEVER committed in the first place — no git op changed the file, no disk lag, no wrong-port. The kernel's in-memory edits were volatile and never persisted. The guard: when moving cells between branches, always verify the COMMITTED state with 'git show branch:file' before trusting your memory of the kernel.
## ruff-f401-cell-split
- ## ruff-f401-cell-split-imports
- When you SPLIT a marimo notebook cell (e.g. splitting a combined startup_form+startup_validation cell into two cells), ruff reports F401 (unused import) false positives on the DEFINING cell's imports. This fires because ruff treats the .py file as a normal Python module and does NOT understand marimo's cross-cell dependency model — in marimo, an import in cell A 'flows' to cell B as a FUNCTION PARAMETER (cell B's signature lists it), not as an in-cell usage. So ruff sees 'import json' in startup_form as unused (startup_form doesn't call json.*) even though startup_validation receives json as a param and uses it heavily.
SYMPTOM: after splitting a marimo cell, 'ruff check notebooks/X.py' reports F401 on imports that ARE used — just by other cells via the parameter-passing model. The assistant burned multiple reasoning turns connecting 'ruff F401' to 'marimo cell parameters' before recognizing this is a false positive, not a real unused import.
ROOT CAUSE: marimo .py files are regular Python from ruff's perspective (NOT .ipynb notebooks, so ruff's notebook-aware logic does not apply). The cross-cell dataflow contract (import -> return tuple -> consuming cell param) is invisible to ruff's static analysis.
FIX (the proper marimo way — wire the dataflow contract per the UNIFIED CELL-DATAFLOW CONTRACT in new-notebook-file-scaffold): (1) the DEFINING cell (startup_form) MUST return all imported names that other cells consume: 'return (json, HTTPError, URLError, Request, urlopen, ...)'; (2) the CONSUMING cell (startup_validation) MUST list those same names as function parameters: 'def startup_validation(json, HTTPError, URLError, Request, urlopen, ...):'. Once the defining cell's return tuple lists the exports, ruff sees them as 'used' (they're in the return statement) and the F401 clears.
ALTERNATIVE FIX (ruff config-level): add a per-file-ignores entry so ruff skips F401 on marimo notebook files: '[tool.ruff.lint.per-file-ignores]' with '"notebooks/*.py" = ["F401"]'. This is appropriate when the notebook uses the app.setup block for shared imports (where imports are GLOBALLY visible without return tuples — see new-notebook-file-scaffold case 1) and you don't want to wire every import through the return/param contract just to satisfy the linter. Note: the build-deep-research-agent repo's [tool.ruff] exclude does NOT list 'notebooks' (only [tool.interrogate] does), so ruff DOES lint notebooks/*.py — making this F401 issue active there.
KEY DISTINCTION: this is NOT a real unused import — the imports ARE used, just across cell boundaries. Do NOT delete the 'unused' imports; wire them through the dataflow contract instead. The import-cell-return section ('a cell that imports a library must explicitly return the imported symbol') is the general principle; THIS section adds the RUFF SYMPTOM that surfaces when you forget to do so, and the per-file-ignores config escape hatch.
Discovered build-deep-research-agent notebook 01 startup_form/startup_validation cell split (07-10): splitting the combined startup cell caused ruff F401 on json, HTTPError, URLError, Request, urlopen — all consumed by startup_validation via cell params.
## server-startup-from-pixi-task
- - STARTING MARIMO VIA A PIXI TASK: when the pixi task is defined as a STRING with embedded subcommands+path (e.g. build-deep-research-agent's 'marimo = "marimo edit notebooks/"'), running 'pixi run marimo <extra-args>' APPENDS the extra args to the TASK STRING, producing a malformed command — e.g. 'pixi run marimo edit --no-token --port 2722 notebooks/00_check.py' expands to 'marimo edit notebooks/ edit --no-token --port 2722 notebooks/00_check.py' (double 'edit', conflicting paths). The server may start but serves the wrong target. BYPASS: invoke the marimo MODULE directly instead of the task alias: 'pixi run python -m marimo edit --no-token --port <p> notebooks/<file>.py' — 'python -m marimo' is not a pixi task so args reach marimo cleanly. DIAGNOSTIC: if 'pixi run marimo <args>' produces unexpected output (two 'edit' commands in the log, serving the dir instead of the file), check whether the task string already embeds subcommands. GENERAL PIXI PRINCIPLE: pixi tasks are string templates, not function calls — to customize a task with embedded args, invoke the binary directly or redefine the task.\n- SINGLE-FILE vs DIRECTORY MODE: 'marimo edit notebooks/00_check.py' (single file) uses a 'single workspace' that can produce 'file_not_found(key)' workspace resolution errors (relative-path resolution quirks or stale browser tabs reconnecting with old session keys). Prefer DIRECTORY mode ('marimo edit notebooks/') which is more robust — the user opens the specific notebook from the browser file list. For 'just the 00 notebook' requests, directory mode still works (marimo-pair targets the session by file path).\n- NOTE: existing env-var-restart commands elsewhere in this skill that use 'pixi run marimo edit notebooks/<f>.py --port <p>' ALSO have this latent double-edit bug when the task prefix is 'marimo edit notebooks/' — use 'pixi run python -m marimo edit ...' in those commands too. Discovered build-deep-research-agent 07-10 (issue-#19 worktree, pixi task TUT-SETUP-012).
## UIElement-same-cell-value-access
- - REFINEMENT (07-10): the ACTUAL error raised when a cell accesses a UIElement's .value in the same cell that created it is a RuntimeError (NOT a CycleError as previously documented), with this EXACT message: 'RuntimeError: Accessing the value of a UIElement in the cell that created it is not allowed. Fix: move the value access to another cell.' This applies to ALL mo.ui.* elements including mo.ui.run_button — a cell that does 'save_env = mo.ui.run_button()' then later 'mo.stop(not save_env.value)' raises this RuntimeError because mo.stop evaluates save_env.value in the same cell. The EXACT message string is greppable: when debugging a marimo cell that references its own widget .value, grep for 'Accessing the value of a UIElement in the cell that created it' to confirm this rule is the cause. Confirmed build-deep-research-agent 00_check.py (07-10): a combined startup cell created base_url_input/model_input/save_env widgets AND called mo.stop(not save_env.value) + read the other .value attrs in the same body → RuntimeError; split into env_form (creates widgets) + env_check (reads .value, mo.stop-gated) fixed it. This is the canonical reason issue #20 split the startup cell into two in the first place — do NOT attempt to re-combine them into one cell; the two-cell split is REQUIRED by marimo's reactive graph, not a stylistic choice.
## pytest-cell-signature-matching
- When writing pytest regression tests that check for the presence of a marimo cell function via source.find() against the notebook .py file, use the OPENING-PAREN form 'def funcname(' — NOT the empty-parens form 'def funcname():'. REASON: a marimo cell that uses reactive inputs from other cells has ALL those inputs as FUNCTION PARAMETERS, and multi-parameter cells span multiple lines: 'def startup_validation(\n HTTPError,\n Path,\n Request,\n ...,\n):' — the string 'def startup_validation():' (empty parens) will NOT match because the parens contain parameters, not nothing. Only zero-parameter cells match the '():' form. CORRECT PATTERN: source.find('def startup_validation(') — match up to and including the opening paren, omit the closing. This is robust to both zero-param and multi-param cells. Example (build-deep-research-agent tests/test_notebook_startup_validation.py): 'startup_index = source.find("def startup_form():")' works because startup_form has NO params; 'validation_index = source.find("def startup_validation(")' must use the open-paren form because startup_validation has 18 params (Request, urlopen, HTTPError, URLError, etc.). DISTINCT from DISK-FILE LAG (that failure: the .py file on disk doesn't have the cell at all — grep returns nothing; THIS failure: the .py file HAS the cell but the test string doesn't account for multi-line params). ALSO DISTINCT from ruff F821: if you ADD parameters to a cell function signature to fix F821 undefined-name errors (symbols used in the body but not declared), the test's '():' match will BREAK as a side effect — update the test to the '(' form in the same change. Discovered build-deep-research-agent notebook 01 startup_validation cell (07-10): after adding Request/urlopen/HTTPError/URLError to the cell params to fix F821, the test source.find('def startup_validation():') stopped matching.
## celoutput-errors-attribute
- CELLOUTPUT .errors ATTRIBUTE (the traceback channel when .data is empty): a marimo CellOutput object (cell.output for a given cell) has attributes: asdict, channel, data, empty, errors, mimetype, stderr, stdin, stdout, timestamp. When a cell has status=exception but cell.output.data is None or empty string, the actual traceback/error message is in cell.output.errors — NOT .data. ALWAYS check .errors (and .stderr, .stdout) when .data is empty for an exception-status cell. The existing diagnostic sections (post-edit ran-error triage, mo-stop-empty-data, duplicate-definition) all say 'read cell.output.data' — but .data is the DISPLAY payload (rendered output), not the ERROR channel. For a stopped/exception cell, .data can be empty while .errors contains the exception traceback. Diagnostic order for any status=exception cell: (1) cell.output.errors → the actual traceback; (2) cell.output.data → the display payload (may be empty); (3) if BOTH empty, check for mo.stop() without a message arg (expected halt, not a real error — see mo-stop-without-message-produces-empty-data). This was discovered the hard way (build-deep-research-agent env_check cell, 07-10): spent multiple turns probing .data (empty), reasoning about mo.stop logic, reproducing in the scratch context — all because the traceback was sitting in .errors the whole time. Do NOT iterate attribute names blindly — .errors is the first thing to check for an exception cell.
## cell-errors-diagnostic-api
- Each marimo cell object exposes a 'cell.errors' attribute (accessed as 'c.errors' when iterating ctx.cells) — a 'list[CellError]' of structured error objects. This is a COMPLEMENTARY diagnostic to 'cell.output.data' (traceback text) and 'cell.status' (idle/exception/marimo-error). Use ALL THREE together when triaging a cell that should have run cleanly: 'c.errors' gives structured error entries, 'cell.output.data' gives the raw traceback string, and 'cell.status' gives the high-level verdict. GOTCHA (build-deep-research-agent 07-10): 'c.errors' can return an EMPTY list (0 errors) even when 'cell.status == exception' — so it is NOT a reliable sole indicator. Observed on a FRESH server (port 2724) after disk-load: env_form was idle (healthy) while env_check was exception but had 0 CellError entries and empty output. Leading hypothesis: the exception status is STALE — set during a prior run/session when there was a real error (e.g. MultipleDefinitionError or a mo.stop interaction), then the error LIST was cleared on a subsequent re-run while the STATUS field was not reset. This means a cell can carry an orphaned exception status that does not reflect its current error-free state. DIAGNOSTIC: when you see 'status=exception, c.errors=[], output=empty', re-run the cell cleanly (ctx.run_cell) and re-check — if status clears, it was stale. OPEN QUESTION (unresolved 07-10): does mo.stop(True) cause cell status to read as 'exception'? The agent was mid-test when the session was cut; do NOT assume mo.stop sets exception status without verifying. This section complements the existing 'post-edit ran-error traceback triage (read cell.output.data...)' guidance by adding the c.errors attribute and the stale-status paradox.
## stale-output-transient-condition
- RESOLVES the dangling 'stale-status-after-rerun' cross-reference (previously only mentioned at line 158 'see stale-status-after-rerun' and line 182 'the stale-status paradox where c.errors returns empty while cell.status reads exception'). A marimo cell's displayed OUTPUT and STATUS reflect the moment it LAST EXECUTED, NOT real-time current state. When a cell shows an error/timeout/danger callout (TimeoutError, connection error, 'not ready', 'exception' status) but the CURRENT conditions look fine — the endpoint responds to a fresh curl/ping, the file exists on disk, env vars are set, a manual repro in a scratch cell works — the cell's output is STALE from a TRANSIENT PAST CONDITION (cold/scaled-down LLM endpoint, momentary network blip, a service that has since warmed up) captured at its last run. It is NOT a logic bug in the cell. DIAGNOSTIC PRIORITY: RE-RUN the cell (ctx.run_cell in code_mode, or the user clicks Run) as the FIRST step before deep-diving into the cell's logic, function params, module globals, or the reactive graph. The re-run executes against CURRENT conditions and the output refreshes — if it now succeeds, the prior failure was transient staleness. Concrete instance (build-deep-research-agent 00_check.py env_check, 07-10): env_check showed a 'not_ready' TimeoutError callout ('endpoint did not respond within 90s') while a fresh curl returned HTTP 200 in 0.9s and a manual scratch-cell repro of the exact same logic returned has_env=True + status=200; the assistant spent 4+ reasoning turns investigating a 'logic bug' (comparing scratch vs cell params, suspecting stale module globals) before realizing env_check's output was a stale TimeoutError from when the endpoint was cold/scaled-down minutes earlier; one re-run showed 'ready=True'. DISTINCT from the other staleness gotchas in this skill: (a) code-mode-edit-flush-on-exit = stale READS within one async-with block before queued edits flush; (b) ctx.cells-stale-after-mutation = stale ctx.cells dict after delete/reorder; (c) StaleCellError = edit_cell REFUSAL across blocks because the cell body drifted; (d) hide_code-verification = hide_code attribute reads False even after a persisted set. THIS section is none of those — it is a cell's OUTPUT reflecting a transient past RUN CONDITION (a cold endpoint that has since warmed), fixed by simply re-running. Diagnostic heuristic: if a scratch/probe cell reproducing the target cell's logic SUCCEEDS while the target cell itself shows failure, suspect stale output and re-run the target cell FIRST.
## mermaid-fenced-block-does-not-render-in-mo-md-use-mo-mermaid
- A ```mermaid fenced code block placed INSIDE mo.md() does NOT render as a diagram in marimo — marimo's markdown renderer does not process the mermaid fenced block into a rendered diagram (it appears as a bare code block, or the block is dropped entirely), so the diagram is invisible/non-rendering even though the markdown renders without error. This is the SAME pattern family as mkdocs-admonition-syntax-does-not-render-in-marimo and mo-callout-needs-mo-md-wrap: markdown syntax that GitHub/Jupyter renders but marimo's mo.md() does NOT. CANONICAL API: mo.mermaid(diagram) — verified marimo 0.23.8 (inspect.signature(mo.mermaid) returns (diagram: str, theme: str | None = None, theme_variables: dict[str,str] | None = None) -> Html). It takes the RAW diagram string (NO ```mermaid fences, NO the literal 'mermaid' header — just the 'graph LR ...' body) and returns an Html object. FIX PATTERNS (two valid approaches, both verified 07-10 via _mime_ output inspection): (A) F-STRING INTERPOLATION into mo.md() — PREFERRED when the diagram sits within a prose narrative. mo.mermaid() returns an Html object, so it CAN be f-string-interpolated into mo.md() (the <marimo-mermaid> web component renders inline within the markdown, with prose before and after). CORRECTION: an earlier version of this section claimed mo.mermaid CANNOT be called inside an mo.md() string (same restructuring constraint as mo.callout) — that was WRONG: mo.mermaid returns Html (interpolatable, same mechanism as mo.as_html(axis) in mo.md f-strings), whereas mo.callout needs a separate mo.md() wrap because its CONTENT arg must already be markdown. Pattern: mo.md(f"""...{mo.mermaid(diagram)}..."""). (B) mo.vstack — alternative when the diagram is a standalone block: mo.vstack([ mo.md(...), mo.mermaid(diagram), mo.md(...) ]). BEFORE (non-rendering): 'def loop_diagram(): mo.md(dedent(""" ### The loop \n\n```mermaid\ngraph LR\n A --> B\n```\n\nSome prose. """))'. AFTER (renders): define the diagram as a raw string, then 'mo.vstack([ mo.md("### The loop"), mo.mermaid(diagram), mo.md(dedent("""Each lap is one **iteration**...""")) ])' as the cell's last top-level expression. GOTCHAS: (1) the diagram string passed to mo.mermaid must NOT include the ```mermaid fences or the literal token 'mermaid' — strip them when extracting from the fenced block; (2) mo.mermaid returns an Html object, so the cell's LAST top-level expression must be the mo.vstack (or the mo.mermaid call itself if it is the only output) — an expression nested inside a statement is discarded (see the last-top-level-expression / expression-in-if-block sections, memory #604); (3) mermaid node labels with quoted text (Think -->|"enough info"| Answer) are fine in the raw diagram string; (4) optional theme / theme_variables kwargs control dark/light theming (marimo auto-picks based on app theme if omitted). SYMPTOM that triggers this: a notebook cell is meant to show a flow diagram / graph / flowchart but renders empty or shows raw 'graph LR' text as a code block — and the cell has status idle (no error). Concrete instance: build-deep-research-agent notebooks/04_workflows.py loop_diagram cell (07-10) wrapped a mermaid 'graph LR' Think/Act/Observe loop inside mo.md(dedent('''...''')); diagram did not render. When converting an existing fenced-block cell to mo.mermaid via code_mode, re-read ctx.cells[target].code first and restructure as a single edit_cell with the mo.vstack body + explicit 'return', then ctx.run_cell to verify the diagram appears.
- GOTCHA (5): an f-string interpolation variant — mo.md(f'''...{mo.mermaid(d)}...''') — RUNS without error (the Python f-string evaluates mo.mermaid(d) to an Html object before mo.md processes the string), but whether the diagram RENDERS VISUALLY via this approach is UNVERIFIED. An agent used this variant in a build-deep-research-agent session (07-10) and declared 'should now render' after only confirming 'cell ran idle, no error' — WITHOUT visual confirmation. mo.vstack remains the RECOMMENDED reliable approach because it passes actual Html objects to marimo's layout renderer directly. VERIFICATION PRINCIPLE: 'cell ran without error' (status=idle) does NOT mean 'diagram rendered' — a cell can run successfully yet still produce non-rendering output (the core issue this section documents). Always confirm the diagram VISUALLY (user checks browser, or screenshot) before declaring the mermaid fix complete. This is the marimo-rendering-specific instance of the general verification-discipline rule (memory #4: heuristics != facts; 'ran without error' is a heuristic, visual rendering is the fact).
## run_cell-triggers-downstream-timeout
- run_cell on a FAST cell can STILL cause the code_mode block to TIME OUT — because run_cell, like set_ui_value (see set-ui-value-downstream-timeout section), triggers marimo's reactive graph to AUTO-RUN downstream cells that depend on the edited cell. If a downstream cell calls the LLM (e.g. ex1_run calls agent(...) which invokes the LLM), the slow downstream execution blocks the run_cell call and the execute-code.sh command times out CLIENT-SIDE — but the edits LAND CORRECTLY. The kernel keeps running the downstream cell regardless. FIX: after editing+running a cell in a notebook where downstream cells call the LLM, DO NOT panic about the timeout — verify cell status (idle/error) in a SEPARATE fresh read (a non-triggering execute-code call that only inspects ctx.cells), rather than re-editing or re-running. The edit is fine; only the downstream LLM execution was slow. This is the run_cell analog of set-ui-value-downstream-timeout (which covers the same pattern for set_ui_value). GENERAL RULE: ANY code_mode mutation (edit_cell, run_cell, set_ui_value) that changes cell state triggers marimo's reactive graph to re-run downstream dependents — if those dependents call the LLM, expect a timeout, and verify edits in a fresh read rather than diagnosing it as an edit failure. Discovered build-deep-research-agent Part 4 ex1_scaffold/ex1_run cells (07-10): edit_cell(ex1_scaffold) + run_cell(ex1_scaffold) timed out because downstream ex1_run auto-ran and called the agent (LLM); both cells verified idle (no errors) in a fresh read.
## batch-reactive-re-run-timeout
- GENERALIZATION of 'set_ui_value triggers reactive re-runs' (that section covers set_ui_value specifically; THIS covers ANY reactive-triggering code_mode op, especially run_cell after edit_cell). When a code_mode 'async with cm.get_context() as ctx:' block contains a BATCH of edit_cell + run_cell pairs (the AGENTS.md-recommended pattern: run_cell after every edit_cell to refresh downstream outputs), EACH run_cell triggers marimo's reactive graph to re-execute ALL downstream dependent cells. In notebooks with LLM-making cells (agent tutorial notebooks, llamabot AgentBot cells, any cell that calls a hosted LLM), each downstream re-run can take 5-30+ seconds per LLM call. A batch of 10+ edit_cell+run_cell pairs feeding into LLM cells can accumulate to MINUTES of reactive re-execution time inside the single async-with block — exceeding the execute-code.sh HTTP timeout (CLIENT-SIDE; the kernel keeps running). SYMPTOM: the code_mode block 'times out' or returns empty/hangs after a long wait; you cannot tell from the client side whether the queued edits FLUSHED before the timeout (the async-with context manager flushes on __aexit__, but if the HTTP client disconnected, the flush may complete server-side later but you cannot observe it). DISTINCTION from mid-batch-exception-discards-flush (THAT is a synchronous exception inside the block → __aexit__ discards the flush deterministically; a CLIENT-SIDE TIMEOUT is different — the block may still be running server-side and may flush eventually, but the agent cannot verify). MITIGATION (pick one): (1) CHUNK the edits — issue 2-3 edit_cell+run_cell pairs per code_mode block instead of 10+, so each block's reactive re-execution completes within the timeout; (2) SEPARATE structure from execution — queue ALL edit_cell calls in one block WITHOUT run_cell (structural changes only, no reactive re-execution), exit, then run cells in small groups in subsequent blocks; (3) TEMPORARILY make downstream LLM cells inert (wrap their body in a conditional or mo.stop) before the batch, do the edits, then restore them — so reactive re-runs are cheap. The key insight: in notebooks with LLM-making cells, the reactive re-execution triggered by run_cell is EXPENSIVE, and batching many run_cells in one block is the timeout trigger — not the edit_cells themselves (edit_cell is structural-only and does NOT trigger reactive re-execution). When restructuring such notebooks, prefer the STOP-REWRITE-RESTART path for major changes, or chunk code_mode edits into small groups. Discovered build-deep-research-agent Part 4 notebook restructure 2026-07-11: a comprehensive code_mode block with many edit_cell+run_cell pairs timed out because LLM-making cells (agent pipeline cells) reactively re-executed on each run_cell.
## hide_code-tutorial-rule
- TUTORIAL NOTEBOOK CELL VISIBILITY RULE (build-deep-research-agent exercise notebooks, distinct from the competition rubric above): the ONLY visible code cells (hide_code=False) are EXERCISE SCAFFOLD cells (exN_scaffold) — the cells where participants fill in their implementation. ALL other code cells — setup, imports, narrative/markdown (mo.md), demo code, helper code, non-exercise logic — must be hide_code=True. Stated as a correction + 'you gotta remember' (07-11): 'most of the other code blocks are hidden, except for the code cells that have an exercise in there.' RELATIONSHIP to the competition rubric: the rubric (~25% visible: method/experiment/training cells) applies to COMPETITION notebooks (alphaxiv-marimo); THIS rule applies to TUTORIAL/EXERCISE notebooks where the visible set is exactly {exercise scaffold cells}. A scaffold cell that became non-exercise (e.g. reduced to a one-liner with no blanks to fill) should ALSO be hidden since it is no longer an exercise scaffold. When authoring or reviewing: enumerate all code cells, confirm each is either (a) an exercise scaffold → hide_code=False, or (b) everything else → hide_code=True. Use ctx.edit_cell(name, code=None, hide_code=True/False) to toggle without disturbing the body.
- EDIT_CELL OMIT-INHERITS GOTCHA (build-deep-research-agent AGENTS.md, 07-11): when you call ctx.edit_cell(name, code='new body...') WITHOUT passing hide_code, the edit SILENTLY INHERITS the cell's prior hide_code state — if the cell was previously hidden (hide_code=True from an earlier edit or UI action), it STAYS hidden even though you changed the code. The edit succeeds, the code updates, but the participant sees nothing. ALWAYS pass hide_code explicitly on edit_cell when cell visibility matters: hide_code=False for scaffold/exercise cells, hide_code=True for infrastructure. This is the WHY behind the hide_code-tutorial-rule's instruction to use edit_cell(name, code=None, hide_code=False) — omitting the parameter is not neutral, it means 'keep whatever was there', which for a previously-hidden scaffold is invisible.
## module-level-constants-dropped-on-rewrite
- FULL-REWRITE FOOTGUN (stop → write .py on disk → restart, the AGENTS.md §8 reliable path): constants written at MODULE LEVEL — i.e. outside ALL three structural block types (@app.cell, @app.function, with app.setup:) — are silently DROPPED from the reactive graph. Marimo does NOT execute bare module-level statements between structural blocks; they are inert decoration in a serialized graph file, NOT plain Python top-level code. THE MISDIRECTION: marimo's static parser still detects cells reference these names and ADDS them as cell-function parameters in the regenerated signatures (def detect_ollama(GEMMA4_MODEL, MIN_RAM_GB, OLLAMA_TAGS_URL):), so the wiring LOOKS correct — but the names are defined NOWHERE (not in app.setup, not in any cell return tuple). At runtime every receiving cell raises NameError, which reads like a stale-kernel / broken-graph bug, NOT a 'constants in the wrong place' bug. FIX: put ALL shared constants in the with app.setup: block (globally visible, no cell param needed — see the 'app.setup block variables are GLOBALS' gotcha). DIAGNOSTIC TELL: NameError on a name that appears in a cell signature but is absent from app.setup AND all cell bodies, right after a full-rewrite. This is the INVERSE of the refactoring gotcha (setup global mistakenly left as a cell param): that one is 'defined in setup AND listed as a param (redundant)'; THIS is 'defined NOWHERE but listed as a param (missing)'. Also: with ... as <var>: targets (with urlopen(req) as resp:) count as cell-level name bindings for the single-definition rule, just like imports/assignments/for-loop targets — two cells both using 'as resp' collide with MultipleDefinitionError. Discovered build-deep-research-agent 00_check.py full-rewrite (07-11): 9 constants written after the app.setup block as module-level assignments; marimo added them as cell params but they NameError'd at runtime.
## ui-radio-value-label-gotcha
- - mo.ui.radio .value ASYMMETRY from mo.ui.dropdown (build-deep-research-agent 07-11): mo.ui.radio(options={'local': 'Local Ollama (gemma4:12b)', 'remote': 'Remote endpoint'}) returned the DISPLAY LABEL string from .value ('Local Ollama (gemma4:12b)') instead of the key ('local'). This CONTRADICTS mo.ui.dropdown(options={key: value}) which returns the dict VALUE from .value (see ui-slider-dropdown-api above). The asymmetry is non-obvious because both use the {key: value/label} dict options format, yet radio returns the LABEL side while dropdown returns the VALUE side. SILENT FAILURE MODE: a check like 'if source.value == "local":' silently NEVER matched (always False) — the wrong code branch executed with NO error, NO exception, just incorrect behavior (writing remote env values when the user selected local). The user saw a broken .env file ('you gotta check my .env file, man. It's not correct') with no indication the radio comparison was the cause. DIAGNOSTIC: always PRINT the element's .value ('print(f"source.value: {source.value}")') to verify what it actually returns BEFORE writing comparison logic — do not assume the format from documentation. FIX: use defensive SUBSTRING matching ('is_local = "local" in source.value.lower()') instead of exact equality ('if source.value == "local":') — this is robust whether .value returns the key or the label. Broader lesson for ALL mo.ui.* elements with {key: label/value} dicts: verify empirically what .value returns for EACH element type (radio, dropdown, checkbox, etc.) before writing equality checks against it, because the return convention is NOT uniform across marimo UI elements.
## celltour-maintenance-verification
- EXISTING-TOUR VERIFICATION (distinct from building a tour fresh): when asked to 'ensure the code tour is up-to-date' or 'check the CellTour' AFTER notebook cells were edited/merged/renamed (not building fresh), the verification is a READ-ONLY cross-reference: (1) grep the on-disk .py for ALL 'def <name>' function declarations to get the current set of valid cell names; (2) read the CellTour steps from the tour cell (grep 'cell_name' or read ctx.cells[tour_cell].code); (3) diff — every cell_name in the tour MUST exist in the def-name set; any name that doesn't is a STALE reference (cell was renamed, renumbered, or deleted). Do NOT compare tour cell_names against ctx.cells keys (auto-generated IDs like Hbol/MJUe) — those are NOT what CellTour resolves; the def function names in the .py file are the authoritative set. COMMON TRIGGER: after a notebook merge (PR merged new cells/renamed exercises), the CellTour cell from a prior session references old cell names that the merge changed (e.g. tour says 'ex2_header' but the cell was renumbered to 'ex1_header' per the exN naming convention). FIX: update the stale cell_name entries in the tour cell to match the current def names. Discovered build-deep-research-agent notebook 04 (07-11): post-merge tour verification required cross-referencing 10 cell_name entries against the disk def names, not the live kernel IDs.
## gotchas
- test_read_only
-
## Delete + create in same code_mode batch leaves stale cells
Batching ctx.delete_cell() + ctx.create_cell() in the SAME
'async with cm.get_context() as ctx:' block can leave STALE CELLS behind.
All ctx.* methods are synchronous and queue operations — the context
manager flushes them on exit. A failed delete (e.g. wrong cell ID/name)
combined with a successful create means the old cell PERSISTS as a ghost,
causing persistent NameError or status='marimo-error' from the stale cell
referencing now-private or deleted variables.
**Fix:** Split into SEPARATE code_mode calls — one context manager for
delete, a second for create. After each batch, iterate
ctx.cells.values() to verify no stale cells remain:
```python
# Step 1: delete (separate context manager)
async with cm.get_context() as ctx:
ctx.delete_cell('old_cell_name')
# Step 2: verify deletion, then create (separate context manager)
async with cm.get_context() as ctx:
stale = [name for name, c in ctx.cells.items()
if 'old_pattern' in c.code]
for name in stale:
ctx.delete_cell(name)
ctx.create_cell(code='...', name='new_cell_name')
ctx.run_cell('new_cell_name')
```
**Diagnostic:** if a cell you thought you deleted still shows
status='marimo-error' or raises NameError for a variable that was
supposed to be gone, iterate ctx.cells.values() — the stale cell is
likely still present from a batched operation that partially failed.
Discovered build-deep-research-agent Part 2 Ex2 observations UI (07-11):
batched delete+create left cell GWlb (old save handler referencing
_ex2_save) alive while the new cell iimj was created alongside it.
## fill-in-blank-cell-signature-contamination
- In marimo notebooks with fill-in-the-blank pedagogical exercises (Network-Analysis-Made-Simple and any tutorial using blank placeholders like ___, ____, G_________ in exercise stubs), pedagogical blanks must NOT leak into the cell function signature. The cell function signature (def _(G, G_________, ___, ____):) declares reactive-graph DEPENDENCIES — marimo resolves each parameter as a variable from an upstream cell. Blank-named parameters (___, ____, _________, G_________, d) are pedagogical placeholders meant for the student's implementation in the function BODY, not real cell dependencies. When they appear in the signature, marimo cannot resolve them and the cell fails.
KEY INSIGHT — solution-import override pattern: when exercise stub functions are overridden by solution imports (from nams.solutions.X import filter_graph), the stub body with blanks NEVER actually executes — the imported solution function is what runs. So blanks in the cell signature are never real dependencies; they are artifacts of the fill-in-the-blank pedagogy that accidentally contaminated the signature.
FIX: strip all blank-named parameters from the cell signature, keeping only real upstream cell dependencies. Example: def _(G, G_________, ___, ____, ___________, d): -> def _(G): — G is the only real cell dependency (defined in an upstream cell); the blanks and even 'd' (which looks like it's used in the body as d[___________] but is actually a fill-in-the-blank for 'data') are pedagogical artifacts.
DIAGNOSTIC: if a cell function signature contains parameter names consisting entirely of underscores (___, ____, _________) or underscore-suffixed identifiers that look like fill-in-the-blank placeholders (G_________, G_), they are pedagogical artifacts, not cell dependencies. Even a 'real-looking' name like 'd' can be a pedagogical blank if the exercise expects the student to fill it in with 'data' — check whether the name is defined in an upstream cell (real dependency) or only appears as a blank to be filled in (pedagogical).
GREP for the pattern across all notebooks: search for cell function signatures containing consecutive underscores in parameter names. Discovered Network-Analysis-Made-Simple 01-io.py (line 430) and 01-bipartite.py (line 173) exercise cells (07-11).
## wasm-html-export
- `marimo export html-wasm` produces a SELF-CONTAINED, INTERACTIVE HTML file that runs Python in the browser via Pyodide (Python-on-WebAssembly). This is DISTINCT from `marimo export html` (a static, non-interactive snapshot with pre-rendered outputs — see competition-sharing-fix). Use `html-wasm` when the reader must INTERACT (sliders, re-run cells, change parameters); use `html` when a frozen rendered view suffices.
- EXPORT COMMAND: `marimo export html-wasm notebook.py -o output_dir --mode run`. The `--mode run` flag auto-executes cells on load. Output is a directory (index.html + assets), hostable on any static server (GitHub Pages, Lektor static dir, iframe embed).
- PLATFORM-SPECIFIC DEPENDENCIES via PEP 508 MARKERS: dependencies that CANNOT run in WASM (anything compiling C at runtime) must be gated with `sys_platform != 'emscripten'` so they load locally but are excluded in the browser. In a `--sandbox` notebook's inline metadata:
```python
# /// script
# dependencies = [
# "numpy",
# "scipy",
# "pymc ; sys_platform != 'emscripten'",
# ]
# ///
```
Then branch at runtime: `import sys; IN_WASM = sys.platform == 'emscripten'`. Use pymc when not IN_WASM, a pure-numpy fallback sampler (Metropolis, emcee) when IN_WASM.
- PRE-INSTALLED in Pyodide (no micropip needed): numpy, scipy, scikit-learn, pandas, matplotlib. Pure-Python packages (no C extension) auto-install via micropip at load. Packages with native/C extensions NOT in the Pyodide pre-built set will FAIL — this is the primary constraint.
- `--sandbox` FLAG: inlines the notebook's dependencies (PEP 723 inline script metadata) into the exported file so it is fully self-contained. Use `--sandbox` for portable single-file sharing; omit when hosting the notebook alongside a managed environment.
- KEY CONSTRAINT (see the pymc-wasm memory): pymc/pytensor, numba, jax (CPU-compile paths), and any package requiring a runtime C toolchain will NOT work in WASM. Plan a pure-Python fallback BEFORE exporting. The cleanest pattern is a single notebook that works BOTH locally (pymc) AND in WASM (numpy sampler) via the IN_WASM branch above.
- DISCOVERED embedding a Bayesian 4PL curve-fit marimo notebook in a Lektor blog post (website, 07-24): single-curve fitting was insufficient (user wanted Bayesian posterior bands), pymc was the local tool but blocked in WASM by pytensor's C compilation, so the notebook branches to a pure-numpy Metropolis sampler under emscripten.
## ipynb-export-tracking
- When a marimo project tracks BOTH `.py` (marimo source) AND `.ipynb` (Jupyter export) files in git, renaming functions/variables/imports in the `.py` files creates STALE REFERENCES in the `.ipynb` exports — the `.ipynb` files import or call function names that no longer exist in the `.py` source. The `.ipynb` files are auto-generated exports (via `marimo export notebook.py --format ipynb` or marimo's export menu) used by bookbuilder/nbconvert/Jupyter pipelines, and they are NOT automatically regenerated when the `.py` source changes.
- DIAGNOSTIC: after any rename/refactor in a marimo `.py` file, grep the `.ipynb` files for the OLD function/variable names: `rg 'old_function_name' --include '*.ipynb'`. If hits are found, the exports are stale. Also check `git status` — if `.ipynb` files appear as tracked but unchanged while `.py` files changed, the exports were NOT regenerated and are stale.
- FIX OPTIONS: (1) REGENERATE — `marimo export notebook.py --format ipynb --output notebook.ipynb` for each notebook (cleanest — produces a fresh export from the current `.py` source). (2) SED-REPLACE — for a small targeted rename, `sed -i 's/old_name/new_name/g' notebook.ipynb` (faster for one or two renames, but fragile if the old name appears in contexts that shouldn't change). (3) DELETE the `.ipynb` from git tracking if the project no longer needs Jupyter compatibility — `git rm --cached *.ipynb` + add to `.gitignore` (prevents the problem permanently for projects that don't need the `.ipynb` format).
- PREVENTION: add a pre-commit hook or CI check that regenerates `.ipynb` exports from `.py` files and fails if the tracked `.ipynb` differs from the freshly-generated one — this catches stale exports before they land in a commit. Alternatively, `.gitignore` the `.ipynb` files and regenerate them only in the bookbuilder/CI pipeline.
- GENERALIZES beyond marimo: any project that tracks AUTO-GENERATED DERIVED ARTIFACTS (`.ipynb` from `.py`, compiled CSS from SCSS, generated docs from source) alongside their source must regenerate or update the derived artifacts when the source changes — a rename in the source silently breaks the derived artifact's imports/references. Always check `git status` for tracked derived files that DIDN'T change when their source did.
- DISCOVERED Network-Analysis-Made-Simple 2026-07-11: 6 marimo notebooks had function renames (e.g. spelling fixes like `unrequitted` -> `unrequited`) in `.py` files; the tracked `.ipynb` exports still imported the old names, creating broken imports in the Jupyter versions used by the bookbuilder pipeline.
## notebook-verification-checklist
- A systematic 6-check verification methodology for marimo notebooks (`.py` files), useful after refactoring, renaming, or batch-editing multiple notebooks. Run these PROGRAMMATICALLY across ALL cells in ALL files — do not eyeball. Each check catches a distinct class of structural defect:
- **C1 — Cell parameter traceability**: every cell function PARAMETER traces to a `return` in some other cell (no orphaned parameters that marimo cannot resolve). Parse each cell's `def _(param1, param2, ...)` signature and confirm each param appears in some other cell's `return (...)` tuple. Catches broken reactive-graph wiring after refactors.
- **C2 — No duplicate returns**: no variable name appears in the `return` tuple of TWO different cells (would cause marimo's MultipleDefinitionError). Parse all return tuples across all cells and check for name collisions. Catches the unique-variable-rule violation at the return level.
- **C3 — Single mo import**: `import marimo as mo` (or `import marimo`) appears in EXACTLY ONE cell per notebook (or in the `with app.setup():` block). Multiple marimo imports cause the multiply-defined collision. Count occurrences across all cells.
- **C4 — Underscore-prefix cell-locality**: all underscore-prefixed names (`_x`, `_plots`, `_dest`, etc.) are LOCAL to their defining cell — they do NOT appear in any OTHER cell's return tuple, parameter list, or body as a cross-cell reference. Catches accidental cross-cell reliance on cell-local names. FALSE-POSITIVE WATCH: a name like `_plots` appearing in two cells is fine if BOTH cells independently define it (e.g. `from nxviz import _plots` in each) — verify the name is DEFINED (assigned/imported) in each cell where it appears, not just referenced. Loop variables (`_dest`) similarly can appear in independent for-loops in separate cells without collision.
- **C5 — Consistent App config**: all files have the same `app = marimo.App(...)` configuration (width, auto_download, etc.). Catches drift where some notebooks have different rendering/export settings. Grep for `marimo.App(` and diff the kwargs.
- **C6 — Ending block**: all files end with the exact `if __name__ == "__main__":\n app.run()` block. Catches notebooks where the ending was clobbered (see main-block-clobber-on-save) or manually edited to a non-standard entry point.
- IMPLEMENTATION: write a Python script using AST parsing (`ast.parse` + `ast.walk` to extract function defs, return tuples, import statements) rather than regex — regex on Python source is fragile for nested function bodies, multi-line returns, and decorator-wrapped cells. The script should output a PASS/FAIL table (file × check) so you can see at a glance which notebooks pass all checks.
- FALSE POSITIVES: automated checking of underscore-prefixed names (C4) can produce false positives when the same `_name` is independently defined in multiple cells (local imports, separate loop variables). Always VERIFY flagged hits by reading the actual cell code — confirm the name is DEFINED in each cell, not just referenced. Do not treat a regex/AST hit as a defect without checking the definition context.
- DISCOVERED Network-Analysis-Made-Simple 2026-07-11: used this 6-check methodology to verify 6 marimo notebooks (378 cells total) after a batch refactor; all passed, with two C4 false positives (`_plots` and `_dest` in `02-airport.py`) confirmed as independent local definitions in separate cells.
## mermaid-subgraphs-dont-render-side-by-side
- Mermaid's 'graph LR' with subgraphs does NOT reliably render subgraphs SIDE-BY-SIDE in marimo — even with left-right direction, subgraphs stack vertically (one on top of the other). This is a Mermaid renderer limitation in marimo's mermaid web component, not a marimo bug. FIX: use mo.hstack([mo.mermaid(diagram1), mo.mermaid(diagram2)]) — TWO separate mo.mermaid() calls (each containing a single graph/flowchart) wrapped in a horizontal stack. Each diagram flows top-to-bottom internally; mo.hstack places them left/right. Do NOT attempt to force side-by-side within a single mermaid diagram using subgraphs + graph LR, invisible links, or branching — it is unreliable. SYMPTOM: user says 'they're not left/right, the two diagrams are still one on top of the other' after you used graph LR with subgraphs. VERIFICATION: have the user confirm the diagrams appear side-by-side in the browser (per the visual-verification principle in the mo-mermaid section — 'ran without error' does not guarantee layout). Discovered build-deep-research-agent notebook 05 architecture comparison (two architecture diagrams: 'custom harness' vs 'opencode + MCP'), 07-11.
## code_mode return SyntaxError
- - CODE_MODE 'return' SYNTAXERROR (distinct from the .py-format return rules above and from #71's mid-cell mo.stop): in code_mode (ctx.edit_cell / ctx.create_cell), a 'return (var,)' statement at the END of the cell — valid in the .py file format where cells are 'def _():' functions — causes SyntaxError: 'return' outside function. Root cause: code_mode compiles the cell code as a STANDALONE MODULE BODY (the compiled __marimo__cell_XXXX.py has NO function wrapper), so 'return' is syntactically invalid at ANY position. This means line 46's claim ('marimo's reactive system only propagates variables that are returned by their defining cell') is TRUE for the .py format but FALSE for code_mode: in code_mode, top-level variables defined in a cell AUTO-PROPAGATE to downstream cells WITHOUT any return statement — marimo's reactive system detects top-level variable definitions in the module body and wires them into the graph. FIX: never use 'return' in code_mode cell code; just define the variable at the top level and reference it in downstream cells. The marimo-pair skill's claim that create_cell 'wraps the code as the body of a cell function — effectively def _(): <code>' describes the CONCEPTUAL parameter-inference model, NOT the actual compilation (which is a standalone module). Discovered Network-Analysis-Made-Simple 07-13: 'return (directed_toggle,)' at end of an edit_cell call failed with SyntaxError; removing the return fixed it and the toggle variable was still visible to the downstream SVG cell. Distinct from #71 (mid-cell return inside if-branch in .py format → mo.stop) and from the import-cell-return rule (line 46, .py format only).
## Export batch test gotchas
- When batch-testing marimo exports (marimo export ipynb/md over many notebooks): (1) use EXIT CODES, not string-matching for 'Error'/'error' — marimo prints 'Warning: Notebook has errors, using top-down order instead of topological' for non-fatal issues, and the word 'errors' triggers false-positive FAIL in a grep-based batch test. A warning with a valid output file is SUCCESS. (2) Duplicate 'import marimo as mo' cells cause MultipleDefinitionError at export time — common when CellTour/wigglystuff cells auto-import mo inside their body alongside a standard mo-provider cell. Standard pattern: ONE cell imports+returns mo; all others receive it via 'def _(mo):'. Discovered Network-Analysis-Made-Simple 2026-07-13.
## creating-new-notebook-write-complete-file-first
- When creating a NEW marimo notebook (not editing an existing running one), write the COMPLETE .py file first using the Write tool — all cells pre-written in marimo's file format with @app.cell decorators, proper parameter lists (inputs), and return statements (outputs) — THEN launch it and verify. This is far more token-efficient than building cells one at a time via the code_mode API (each cell = one HTTP round-trip + one tool call).
WORKFLOW:
1. Write the complete notebook .py file using the Write tool:
- Include the marimo header (import marimo; app = marimo.App())
- Write every cell as @app.cell with a unique function name
- Cell function PARAMETERS = inputs (variables received from other cells)
- Cell function RETURN = outputs (variables this cell provides to other cells)
- For shared imports/constants, use 'with app.setup():' (see unique-variable-rule + app.setup patterns)
- End with app.run()
2. Launch the notebook (background, directory mode, no-token):
- Outside a project: nohup uvx marimo@latest edit <dir>/ --no-token --sandbox &>/tmp/marimo.log &
- Inside a pixi project: nohup pixi run python -m marimo edit <dir>/ --no-sandbox --no-token &>/tmp/marimo.log &
- Use DIRECTORY mode (not single-file) to avoid file_not_found websocket errors (see finding-marimo.md gotcha #3)
3. Wait 2s, then discover the server: bash ~/.agents/skills/marimo-pair/scripts/discover-servers.sh
4. Verify by triggering the reactive cascade — run the imports/setup cell via ctx.run_cell() (code_mode), which cascades through all dependent cells. A successful cascade (all cells re-execute without error) confirms the graph is wired correctly and all variables are threaded.
5. Fix any issues interactively via code_mode API (rename cells, add CellTour, adjust styling)
6. Review with a subagent that also connects via marimo-pair
WHY NOT cell-by-cell: building a 20-cell notebook via code_mode API costs 20+ HTTP round-trips and tool calls just for cell creation, before any execution/verification. Writing the complete .py file costs ONE Write tool call, then launch + verify. The tradeoff: you may need to debug variable-threading issues after launch (returns/params must be exactly right), but this is still cheaper than 20+ sequential API calls.
POST-LAUNCH VERIFICATION GOTCHA: if ctx.run_cell() on an upstream cell cascades but some downstream cells still show NameError for variables that ARE returned upstream, check (a) the return statement includes the variable, (b) the downstream cell's parameter list includes it, (c) no duplicate variable definitions across cells (marimo's unique-name rule — see unique-variable-rule / MultipleDefinitionError diagnostic), and (d) import cells actually RETURN the imported symbols (see import-cell-return). The reactive graph threads returned variables to cells that declare them as parameters — a missing param, a missing return, or a duplicate definition breaks the chain silently and the NameError looks like a graph-wiring bug when it is actually one of these three. Discovered 07-14 building multi-stage xarray DLS nanoparticle characterization notebooks in the skills repo.
## Philosophy
- - **Operate within the reactive graph, NOT via `marimo run`.** `marimo run` is SCRIPT MODE — a different execution model that treats the notebook as a linear script. The user pairs with the INTERACTIVE reactive graph (cells, reactive cascades, live browser view). Use the marimo-pair toolchain (discover-servers + execute-code + code_mode API: ctx.edit_cell, ctx.run_cell, ctx.create_cell) for ALL notebook work: reading cells, fixing errors, running cells, verifying state. Do NOT default to `marimo run` to execute or debug a notebook — its errors (e.g. duplicate variable names across cells) signal script-mode constraints, not problems with the notebook in the reactive graph. Stated 07-14: 'no no, don't run as script, you should really lean into operating within the reactive graph.'
## requires-python-change-needs-server-restart
- ## Changing requires-python in PEP 723 metadata requires killing and relaunching the marimo server
The marimo sandbox (uvx marimo edit --sandbox) determines its Python version at SERVER STARTUP TIME from the notebook's PEP 723 requires-python field. Editing the requires-python field while the server is running does NOT change the active sandbox's Python — the running kernel continues on the Python version it booted with, and file edits to the PEP 723 block are also clobbered by marimo's next save (same rule as any file edit; see the "Do NOT hand-edit" guidance in the PEP 723 dependencies section).
**Procedure to change requires-python:**
1. Kill the running marimo server (the session must be down so file edits are safe — see SKILL.md "NEVER Edit/Write the notebook file while a session is running").
2. Edit the PEP 723 # /// script block's requires-python line directly in the .py file (e.g. >=3.11 to >=3.12).
3. Relaunch: uvx marimo edit --sandbox <notebook>. The new sandbox will resolve and use a Python matching the updated requires-python.
**When to do this:** when a dependency requires a newer Python than the current requires-python minimum (e.g. PyMC 6.1.0 requires >=3.12 but the notebook says >=3.11), or when aligning the notebook with a user preference for a specific Python version baseline.
Discovered 2026-07-14 (skills repo, ELISA marimo notebook): PyMC 6.1.0 required >=3.12 but the PEP 723 block said >=3.11; the user directed "you may need to kill server to reset python version" after the assistant reasoned through the conflict.
## topology-stale-node-dangling-reference
- - SYMPTOM: KeyError('<stale_cell_id>') raised inside marimo's transitive_closure when you call edit_cell/create_cell/run_cell on ANY cell whose descendant or predecessor computation traverses a stale node (e.g. KeyError: 'mMEX' when editing stage6_features whose graph walk hits mMEX). The error originates in marimo's graph-traversal, not in your cell code. ROOT CAUSE: a PRIOR cell edit (edit_cell or create_cell) FAILED partway through and left a stale cell ID in the topology's internal edge dicts — the ID exists in topology._children and topology._parents but was never added to (or was removed from) graph.cells, leaving orphaned edges that transitive_closure chokes on.
DIAGNOSIS (do this before STOP-REWRITE-RESTART): access the live graph + topology via code_mode. Check whether the KeyError'd ID is ABSENT from graph.cells but PRESENT in the topology edge dicts: 'stale_id in graph.cells' returns False; 'stale_id in ctx._graph.topology._children' returns True (often mapping to an empty set()); 'stale_id in ctx._graph.topology._parents' returns True (e.g. mapping to {'ecfG'}). This asymmetry (in edges, not in cells) confirms a dangling-node corruption.
IN-SESSION FIX (preserves all live kernel state — try this FIRST, before the nuclear STOP-REWRITE-RESTART): directly prune the stale ID from BOTH topology edge dicts — 'del topology._children[stale_id]; del topology._parents[stale_id]'. ALSO clean any REVERSE reference: if topology._parents[stale_id] listed a parent P (e.g. 'ecfG'), check whether topology._children[P] still contains the stale ID and discard it (topology._children[P].discard(stale_id)). Then VERIFY consistency: assert len(graph.cells) == len(topology._children) == len(topology._parents) (all three counts must match). Retry the originally-failing edit — it now succeeds because transitive_closure no longer encounters the orphaned node.
DISTINCT from reactive-graph-registration-corruption (the sibling section above): THAT corruption is 'a cell runs fine but its output variable never enters the shared namespace' (registration-level), requiring STOP-REWRITE-RESTART as the definitive fix. THIS corruption is 'a stale cell ID lingers in topology edges after a failed edit' (edge-level), fixable by a surgical in-session prune. Try the prune FIRST (cheap, preserves state); fall back to STOP-REWRITE-RESTART only if the prune does not resolve the KeyError or if graph.cells itself is inconsistent. Discovered 2026-07-14: KeyError 'mMEX' from a prior failed edit; pruned from topology._children + topology._parents, counts matched 38/38/38, the retried edit succeeded immediately.
## troubleshooting-multi-session-divergence
- ### Multiple sessions of the same file: code_mode probes the WRONG session
**Symptom.** You run a `code_mode` probe (e.g. iterating `ctx.cells` to find a cell by name) and it returns MISSING — the cell does not exist in the kernel. But the USER says "it has <cell_name> as a cell" (they can see it in their browser). You reason the kernel is stale or the cell was never loaded, and start planning a recovery. The contradiction: the user sees the cell, your probe says it does not exist.
**Root cause.** The same `.py` notebook file can have MULTIPLE open sessions on the same server (e.g. `s_yhs9hh` and `s_vwolwc` for `02-paths.py`). Each session has an INDEPENDENT in-memory state — they loaded the file at different points in time, or one was edited via the browser / a prior `code_mode` call while the other was not. The `--session` flag on `execute-code.sh` routes your HTTP request to ONE session's kernel; `cm.get_context()` returns the context for THAT session. If you probed session A but the user's browser is on session B, your probe reflects A's state, not what the user sees.
**Critical fact — cids are positional/stable but content DIVERGES across sessions.** The same CellId (e.g. `RGSE`) maps to DIFFERENT cell names/code in different sessions: in one session `RGSE` = `bfs_algorithm_reveal` (old), in another `RGSE` = `bfs_animation` (edited). So you cannot assume a cid means the same thing across sessions of the same file. Each session is an independent fork of the notebook state.
**Diagnostic — enumerate ALL sessions before concluding a cell is missing.** Do NOT reason from a single probe. When the user says a cell exists but your `code_mode` probe says it does not:
1. List ALL sessions on the server (the server's session-list endpoint, or `bash scripts/discover-servers.sh` which shows sessions per server).
2. For EACH session of the target file, probe independently: `execute-code.sh --session <id> -c "..."` — count cells, list cell names, check for the expected cell.
3. Identify which session the USER is on (the one whose state matches the user's description) vs. which session is STALE (loaded before your edits).
4. Only THEN decide on recovery: sync the stale session(s) via `edit_cell`, or conclude the user's session already has what it needs.
**The stale session is a clobber bomb.** A session that loaded the file BEFORE your disk edits holds OLD in-memory state. If that session saves (autosave or any user action in its tab), it overwrites the disk file — destroying your edits. This is the multi-session variant of the direct-file-edit recovery problem: not only must you sync the session the user is on, you must also check whether OTHER sessions of the same file are stale and could clobber the disk on their next save.
**Fix.** Once you identify the divergent sessions:
- For the session the user is ON: sync it to disk via `ctx.edit_cell` (push any disk edits that session is missing — see the Recovery procedure in the NEVER-Edit guard rail).
- For STALE sessions the user is NOT on: they are still clobber bombs. Either sync them too (via `--session <stale_id>` + `edit_cell`), or ask the user to close them. Do NOT leave a stale session open — its next save destroys your work.
Discovered 07-14 (Network-Analysis-Made-Simple): `02-paths.py` had two open sessions (`s_yhs9hh` with 36 OLD cells, no `bfs_animation`; `s_vwolwc` with `bfs_animation` from the first task but not the second). The agent probed `s_yhs9hh` first, concluded the cell was missing, and spent multiple reasoning turns on kernel-staleness hypotheses before discovering the second session the user was actually on.
## run-cells-after-edit-create
- ALWAYS run marimo cells after editing (ctx.edit_cell) or creating (ctx.create_cell), AND verify downstream cells re-executed WITHOUT ERROR. This is a NON-NEGOTIABLE workflow discipline, not a hint: create_cell/edit_cell are STRUCTURAL-only (they do NOT auto-execute), so an un-run edited/created cell leaves the notebook with STALE output the user sees immediately in their live browser. Recurring user correction (04-07 AGENTS.md rule + 07-15 'you should run the cells always, I still see cells that are stale and need to be run' — recurred ~3 months apart because it was only weakly in prose). Post-edit workflow: (1) ctx.run_cell(cid) the changed cell; (2) verify its downstream dependents re-ran cleanly — run_cell triggers marimo's reactive cascade, but a downstream cell that ERRORED (status=='marimo-error') looks 'stale' to the user even though it 'ran'; (3) if any downstream cell errored, fix it before declaring the edit done. Leaving stale or errored cells = an incomplete edit. This is the marimo-pair analog of verification discipline (don't trust 'I edited the cell' without running it).
UNRELIABLE-ATTRIBUTE DIAGNOSTIC (the token-sink): when the user reports stale/errored cells, do NOT probe cell.run_result or cell.stale to find them — both are UNRELIABLE. cell.run_result returns a falsy/None-ish value that makes EVERY cell look 'NOT RUN' (clearly wrong for cells that have run), and cell.stale reports False for cells that ERRORED on their last run (an errored cell is not 'stale' — it ran, but errored). The RELIABLE way to find the cells the user sees as stale: iterate ctx.cells.values() and check each cell's STATUS attribute for the value 'marimo-error'. Cells with status=='marimo-error' are exactly the ones needing attention — then read cell.output.data for the cause (MultipleDefinitionError, NameError, etc.). Note: ctx.cells keys are internal CellIds (e.g. 'Hbol'), not names, so when reporting WHICH cell errored, read cell.name or cell.code to identify it, not the dict key. Discovered build-deep-research-agent (07-15): the agent burned several reasoning turns probing cell.run_result (all false), then cell.stale (all false), before checking status and finding 2 cells with status='marimo-error' — the actual cause of the user's 'stale' complaint. The exact status-attribute name may vary by marimo version (status vs _status vs run_result_status) — probe dir(cell) if the attribute isn't found, but the VALUE to match is always 'marimo-error'. Sibling of stale-cell-error (that covers StaleCellError RAISED BY edit_cell; THIS covers DETECTING already-errored cells the user sees as stale).
## edit_cell code must NOT include return statements
- **CRITICAL: the code passed to ctx.edit_cell/create_cell must NOT include the 'return (...)' tuple from the .py file format.** edit_cell compiles code as a standalone MODULE body first (where top-level 'return' is invalid), causing SyntaxError: 'return' outside function. The .py file format uses 'return Index, mo, np, xr' to declare cell outputs, but the code_mode API does NOT want them — marimo auto-detects assigned/defined names as outputs. When copying cell code from the .py source for an edit_cell call, STRIP the 'return ...' line first. Confirmed 07-15 (scipy_2026_causal_inference_tutorial): hit twice in one session.
## batch-hide_code technique
- Added 'Batch-setting hide_code without changing code' technique to notebook-improvements.md: ctx.edit_cell(target, code=None, hide_code=True) sets hide_code while preserving existing code. code=None IS valid for config-only changes (hide_code, disabled); the 'requires full code= body' gotcha applies to name= changes specifically. Discovered xarray NDIndex tutorial 07-15.
## stale-kernel-after-package-install-version-downgrade
- **GOTCHA — ctx.packages.add does NOT reliably update already-imported modules in the running kernel.** When a package that is ALREADY IMPORTED in the running kernel gets a VERSION CHANGE (especially a downgrade) via ctx.packages.add, the on-disk packages update correctly but the running kernel keeps the OLD version in memory. Already-imported modules are NOT reloaded. SYMPTOM: an ImportError or version-conflict error whose version number CONTRADICTS the install output. VERIFIED example (07-15, marimo sandbox): numpy was downgraded 2.5.1->2.4.6 on disk (confirmed by install log), but np.__version__ in the running kernel was still '2.5.1'. numba (installed against numpy <=2.4) detected the stale in-memory numpy 2.5 and raised 'ImportError: Numba needs NumPy 2.4 or less. Got NumPy 2.5.' DIAGNOSTIC: if an import error references a version that contradicts the install output, check the module's __version__ in the running kernel vs the install log — they diverge when a DOWNGRADED package was already imported. FIXES: (1) Restart the kernel via the marimo browser UI (marimo-pair code_mode exposes NO kernel-restart API — ctx has no restart_kernel method). (2) Stop + restart the marimo server and reconnect. (3) Bypass the conflicting backend entirely — e.g. if numba conflicts with numpy version, use PyMC's nuts_sampler='numpyro' (JAX backend) instead of the default numba linker. Do NOT use importlib.reload(numpy) — it is fragile because dependent libraries (numba) cache against the old version at import time. The marimo-pair SKILL.md (line 548-549) claims ctx.packages.add 'handles kernel restarts and dependency resolution correctly' — this is WRONG for version downgrades of already-imported packages. GENERAL PRINCIPLE: package version changes do NOT propagate to a running kernel's already-imported modules regardless of the sandbox/install mechanism. Discovered 07-15 in a scientific Bayesian-modeling marimo notebook with PyMC/numba.
## loop-variable-cell-level-definition-collision
- Variables assigned INSIDE a for loop are treated as CELL-LEVEL definitions by marimo's AST analysis, triggering MultipleDefinitionError collisions with other cells. Example (07-15): a cell build_trial_data has 'for trial_i in range(n_trials): baseline = trial_amplitudes[trial_i] * np.sin(...)' — marimo detects 'baseline' as DEFINED in that cell (the last-assigned loop value leaks to cell scope). If another cell (e.g. epoch_demo) also defines 'baseline' (even as a pm.Deterministic), the dry-run validation raises 'baseline is already defined in cell Xref'. FIX: rename one instance so all cell-level names are globally unique across ALL cells. marimo's single-definition rule applies to ALL assignments including loop bodies and conditional branches, not just top-level statements. This extends the unique-names enforcement (memory #72) with the loop-variable gotcha it didn't cover.
## restart_session-endpoint-frontend-dependent
- marimo's /api/kernel/restart_session HTTP endpoint (same router prefix as /api/kernel/execute) only CLOSES the session — the actual restart depends on the FRONTEND (browser) doing a full reload/reconnect. The handler's own comment states: 'This just closes the session, and the frontend will do a full reload, which will restart the session.' Calling it via HTTP from a headless agent (execute-code.sh, curl, any non-browser caller) leaves the session DEAD because there is no browser to drive the reload. Do NOT call /restart_session to fix a stale kernel from a headless context. When the kernel is genuinely stale (e.g., loaded a numpy/numba version that conflicts with the updated on-disk env) and MUST be restarted: ask the user to restart via the marimo UI (the user does NOT want kernel restarts as a default — recovery via edit_cell is preferred per marimo-pair skill), or bypass the conflict at the application level (e.g., pm.sample(nuts_sampler='numpyro') to bypass numba via JAX). Source-verified 2026-07-15.
## packages-add-restart-reloads-from-disk
- SIBLING GOTCHA of stale-kernel-after-package-install-version-downgrade (that section covers packages.add NOT restarting; THIS covers it DO restart). ctx.packages.add for a GENUINELY MISSING package (plotly, scipy, matplotlib — confirmed missing via a prior `import` ImportError probe) triggers a kernel RESTART, and on restart marimo reloads the notebook from DISK (the saved .py file), NOT from the in-memory cell state. CONSEQUENCE: any code_mode edits (create_cell/edit_cell) pushed BEFORE the restart that have not yet been flushed to disk are at RISK of being dropped — the post-restart kernel reflects the on-disk file, which may be the OLD code. VERIFIED 2026-07-15 (skills project, xarray-tutorial notebook spruce-up): ctx.packages.add for plotly/scipy/matplotlib restarted the kernel; "The kernel restarted and reloaded the OLD notebook code (12 cells), my edits not yet applied." FIX — ORDERING RULE: run ctx.packages.add FIRST, before pushing ANY code_mode edits. Let the install + restart settle on the clean/old on-disk state, THEN push all cell edits (edit_cell/create_cell) into the fresh kernel. This sidesteps the race entirely because there are no unsaved in-memory edits to lose when the restart fires. If you have ALREADY pushed a batch of code_mode edits and THEN realize a package is needed, do NOT call packages.add (it may restart and drop them): either re-push the edits after the restart, or persist the notebook to disk first so the restart reloads your new code. TOGETHER WITH the sibling section, packages.add restart behavior is VERSION/CONTEXT-DEPENDENT — sometimes it restarts (this section: reloads disk, drops in-memory edits), sometimes it does not (sibling: stale in-memory modules). Diagnose which by checking after the call: does the kernel state reflect a reload (ctx.cells count/codes reverted to disk) or the prior in-memory state (stale version)? DISTINCT from the lint/direct-edit clobber (kernel writes its in-memory state TO disk on save): THIS is the REVERSE — a restart reloading DISK into the kernel and dropping in-memory code_mode edits.
## anywidget-esm-json-int-key-coercion
- When building anywidget widgets whose _esm consumes Python data serialized to JSON (especially networkx/graph node IDs, or any dict with INTEGER keys): json.dumps silently coerces INTEGER dict keys to STRINGS (JSON object keys must be strings: json.dumps({256: 0}) == '{"256": 0}'), but list/array-element VALUES keep their type (json.dumps([{"id": 256}]) keeps id as JSON number 256). In the ESM this creates a NUMBER-vs-STRING type mismatch: step.dist[n.id] does a number-key lookup (256) against a string-keyed object ('256') -> returns undefined, silently breaking the widget with NO error thrown. SYMPTOM SIGNATURE: an anywidget works correctly for STRING-id toy graph data (node ids 'S','A','C') but silently breaks for INT-id real data (node ids 256, 257) — lookups in dist/adj maps return undefined, highlights/animation steps don't fire. DIAGNOSTIC: the asymmetry is that DICT KEYS cross the boundary as strings while LIST VALUES keep their JSON type, so two parts of the same payload arrive in ESM with inconsistent id types. FIX (preferred, Python side): normalize ALL node IDs to str() BEFORE serialization — in BOTH the adjacency/distance dict keys (adj = {str(n): [...]}) AND the node objects' id field (nodes.append({"id": str(n), ...})) — so the ESM consistently compares strings to strings everywhere. FIX (alternative, ESM side): coerce String(n.id) at every dict lookup in the ESM (step.dist[String(n.id)]). The Python-side normalization is preferred because it fixes the root inconsistency rather than patching every ESM lookup site. Recurs across any Python+ESM anywidget project passing graph/dict-with-int-key data (Network-Analysis-Made-Simple BFS animation widget; applies to any KG/graph visualization anywidget). Distinct from the emoji-surrogate-pair ESM escape gotcha (that is an ENCODING issue in ESM string literals; THIS is a TYPE-COERCION issue at the JSON serialization boundary).
## widget-reactivity-verification
- When you need to verify a marimo widget's REACTIVITY end-to-end (e.g. a dropdown that triggers a pipeline function on change) but cannot easily drive the UI element from a scratch/execute-code cell, VERIFY THE UNDERLYING PIPELINE FUNCTION DIRECTLY instead of trying to simulate the UI value change. Concretely: if a dropdown cell calls build_real_anim(choice) on change, call build_real_anim(test_value) in the scratch cell and confirm it produces valid output shapes — if the data function works, the widget cell will work when the user changes the dropdown. WHY direct UI simulation is hard: ctx.set_ui_value (marimo code_mode) requires the actual UI element object, which is cell-local and may not be reachable from an isolated scratch/execute-code context; the agent spent reasoning turns on this before pivoting. The pivot to 'call the pipeline function directly' is the token-efficient verification — it exercises the entire real-data path (load → transform → output) without needing to touch the widget. Pair with the toy-vs-real asymmetry lesson: ALWAYS test the pipeline function with REAL-data inputs (real sociopatterns graph, not just the toy graph), because toy graphs use string node IDs and hide type-coercion bugs that real numpy-int64-ID graphs expose. Discovered Network-Analysis-Made-Simple BFS-animation widget (marimo anywidget + networkx).
## wasm-export-blog-embedding
- To embed an INTERACTIVE marimo notebook in a static blog post (Lektor .lr body, or any HTML/Markdown page), use: marimo export html-wasm notebook.py -o output_dir --mode run. This produces a self-contained directory that runs Python in the browser via Pyodide (WebAssembly) — sliders, plots, and reactive linked cells all work client-side with NO server. Host the output directory alongside the blog post (e.g. in the same content folder) and IFRAME it from the post body. Pyodide ships with numpy, scipy, and matplotlib pre-installed, so scientific/data-science demos (curve fitting, distributions, etc.) run without bundling extra packages. This is the MARIMO PATH for interactive blog visualizations; the VANILLA-JS PATH (hand-written <script> demos with drag handlers, fitted curves, etc.) is the alternative when you need full UI control or want zero Python runtime overhead. Two recurring vanilla-JS demo gotchas (Bayesian outlier-exclusion blog post, website 07-24): (1) TEMPORAL DEAD ZONE — declaring 'const n = arr.length' AFTER it is used in a for-loop throws a silent ReferenceError and nothing renders; declare all const/let at the top of the function. (2) DRAG-DESTROY RACE — if a drag handler's render() rebuilds the DOM from scratch (innerHTML / replaceChildren), it destroys the element the user is mid-drag on, breaking the interaction; use an init/update architecture where the drag target persists across renders and only its attributes (cx/cy, transform) are mutated.