Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

Back to skills

Jvm Ml Inference

ASecurity

Engineering CPU and accelerator-backed ML inference from JVM applications: choosing in-process versus remote serving, bounding native sessions and predictors, coordinating engine and request parallelism, batching under a latency deadline, reusing direct buffers, warming deployments and diagnosing native memory outside NMT. Use when DJL, ONNX Runtime or another native inference engine loses throughput as concurrency rises, leaks RSS, overloads a model pool or needs graceful degradation. Model ...

2 stars
0 votes
0 copies
0 views
Added 9/19/2026
developmentgojavatestingapidocumentation

Works with

api

Security Analysis

A100/100

Scanned 9/19/2026

Install to Claude Code

$npx -y skills add robsonkades/agent-skills --skill jvm-ml-inference --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Jvm Ml Inference?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Jvm Ml Inference
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/robsonkades-jvm-ml-inference/badge)](https://www.skillsdirectory.com/skills/robsonkades-jvm-ml-inference)

More formats (shields.io, HTML) on the badges page.

Download Zip
Files
SKILL.md
---
name: jvm-ml-inference
description: >
  Engineering CPU and accelerator-backed ML inference from JVM applications: choosing in-process
  versus remote serving, bounding native sessions and predictors, coordinating engine and request
  parallelism, batching under a latency deadline, reusing direct buffers, warming deployments and
  diagnosing native memory outside NMT. Use when DJL, ONNX Runtime or another native inference
  engine loses throughput as concurrency rises, leaks RSS, overloads a model pool or needs graceful
  degradation. Model quality and training pipelines are outside scope.
---

# JVM ML Inference

## Purpose

Treat inference as a bounded native resource and queueing system. More request threads do not create
CPU, accelerator streams, native sessions or memory bandwidth; they can oversubscribe the engine and
worsen both throughput and tail latency.

## Workload contract

Use the request and existing artifacts to identify the affected inference path and decision. For
capacity or serving design, record model/version, engine/provider/version, target devices,
input shapes and batch distribution, pre/post-processing, in-process or remote boundary,
session/predictor ownership, engine thread
settings, admitted concurrency/queue, warm-up state, latency SLO, useful throughput, heap/RSS/native
memory and failure/fallback semantics. For a focused correction, collect the subset needed to
establish its contract; missing unrelated measurements need not block it.

Inspect build toolchains, compiler release, runtime image, resolved Java/native artifacts and
provider/device support. This skill prescribes no universal JDK baseline; engine requirements
and virtual-thread guidance are version-specific (virtual threads are final in JDK 21).
Do not upgrade Java or the engine merely to apply an example from current documentation.

## Workflow

1. When choosing or reconsidering the serving boundary, compare in-process versus remote serving
   from latency, isolation, scaling, model cadence, accelerator sharing, failure domain and
   operational ownership. Preserve an adequate existing boundary.
2. For native lifecycle changes or memory diagnosis, inventory the affected resources and lifetimes.
   Bound sessions, predictors, arenas, tensors and direct buffers; close resources that expose
   ownership/close contracts and bound retained storage where reclamation is GC-managed.
   Reuse only after actual native completion.
3. For concurrency tuning, use the relevant axes of outer requests, session count,
   intra-op/inter-op threads and device streams. Start from current settings and a measured
   bottleneck; measure absolute goodput and tail latency, not speedup alone.
4. If batching is used, bound size/storage and maximum wait, and dispatch before the earliest
   member deadline minus the execution/remaining-work budget. Test sparse and burst traffic;
   full-batch throughput is not a latency policy.
5. For deployment/readiness changes, warm the JVM code path and the model/engine separately,
   then gate readiness on a representative successful inference rather than model-file load.
6. When changing admission or failure behavior, test overload in an isolated or authorized load
   environment. Use bounded admission, deadline-aware rejection/cancellation and an explicit
   fallback, if required by the contract, whose quality is also verified.

## Decision rules

- Pool only resources documented as non-thread-safe or expensive to create. Pool size must match a
  measured useful concurrency limit, not request concurrency. A documented shareable session
  with bounded admission may already suffice; non-thread-safe resources can also be confined
  or serialized without a pool.
- Reuse direct, native-order buffers when the API permits, with exclusive ownership across
  filling, native execution and result consumption. Moving allocation from heap to direct
  memory inside the hot path does not remove allocation or guarantee zero device copies.
- Native CPU work can retain a virtual-thread carrier and does not gain throughput from virtual
  threads. Isolate/admit it with a bounded executor when necessary.
- NMT excludes many third-party native allocations. Compare process/cgroup RSS with NMT categories
  and application counters for live sessions/tensors; use native profilers where required.
  RSS growth alone does not identify a leaking owner, and RSS minus NMT is not a leak-size metric.
- `jdk.VirtualThreadPinned` absence cannot clear CPU-bound time inside native code; the event needs a
  relevant park/block path to become visible.
- A cancelled Java future may not stop native computation. Define abandonment, late completion and
  resource reclamation explicitly.

## Evidence output

For a focused review, return the finding, supporting contract/evidence, consequence and smallest
correction or check. When measurements are missing, keep causal diagnoses and sizing conditional
and name the evidence that would distinguish the hypotheses. For experiments, separate measured
result from analytical ceiling, pin environment and raw output, and compare the same metrics
after a change. Report what was verified and what remains unknown; an adequate existing design
is a valid outcome.

## References

- [Native resources, parallelism and batching](references/native-resources-and-batching.md) — read
  when reviewing native ownership, sizing pools, engine threads, buffers or a dynamic batcher;
  includes versioned library sources and diagnostic coverage limits.
- Use `jni-and-ffm`, `off-heap-memory`, `concurrency-limiting-and-bulkheads` and
  `load-testing-advanced` for their owning mechanisms.

Attribution

robsonkadesrobsonkades
View sourceMore from robsonkades →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Browser Extension Developer

Use this skill when developing or maintaining browser extension code in the `browser/` directory, including Chrome/Firefox/Edge compatibility, content scripts, background scripts, or i18n updates.

281612 votes

Seo Optimizer

SEO optimization with keyword analysis, readability assessment, technical validation, content quality. Use for search rankings, blog posts, content audits, or encountering keyword density, readability scores, meta tags, schema markup errors.

2132 votes

Google Official Seo Guide

Official Google SEO guide covering search optimization, best practices, Search Console, crawling, indexing, and improving website search visibility based on official Google documentation

1862 votes

Tanstack Start

Build a full-stack TanStack Start app on Cloudflare Workers from scratch — SSR, file-based routing, server functions, D1+Drizzle, better-auth, Tailwind v4+shadcn/ui. Use whenever the user mentions TanStack Start, asks to scaffold a full-stack Cloudflare app with SSR, wants an SSR dashboard, or asks for a React 19 + Cloudflare Workers app with file-based routing and server functions — even if they don't name TanStack Start specifically. No template repo — Claude generates every file fresh per ...

9881 votes

Pentest

PTES-aligned adversarial security audit for backend, frontend, and mobile applications. Produces a CVSS-scored Hacker Report with verified PoCs and phased remediation.

5491 votes
View all in development →