
Claude Skills by hiyenwong
github.com/hiyenwongFramework for scaling agentic AI through system-level harness design — context governance, trustworthy memory, dynamic skill routing, and verification — treating the structured execution layer around foundation models as a first-class design object.
Integration testing patterns for autonomous agent frameworks — mocking LLM routers, verifying tool-use loops, contract validation, and fallback chains. Applies to super_factory and similar spec-driven agent architectures.
Design and implement memory-augmented AI agents using modular architecture (extraction, update, retrieval, response). Inspired by MemFactory (arxiv:2603.29493) - unified training/inference framework for agent memory with RL-driven policy optimization (GRPO). Use when building long-term AI agents, memory management systems, or implementing Memory-R1/RMM/MemAgent paradigms. Keywords: agent memory, memory-augmented LLM, MemFactory, Memory-R1, memory lifecycle, GRPO, memory extraction, memory ret...
Memory forgetting techniques for autonomous AI agents - adaptive budgeted forgetting, relevance-guided scoring, and bounded optimization for managing long-horizon agent memory systems. Prevents temporal decay and false memory propagation. Source: arXiv:2604.02280. Activation: agent memory, memory forgetting, relevance scoring, memory budget, false memory prevention, adaptive forgetting.
Agent Memory Characterization and System Implications - LLM代理内存系统的首次系统性特性分析和10项系统建议
Agent² RL-Bench: Benchmark for evaluating agentic RL post-training where LLM agents autonomously design, implement, and run complete RL pipelines. Use when evaluating LLM agent capabilities for reinforcement learning engineering, RL pipeline automation, or agentic model alignment.
Automated research pipeline for searching papers, extracting patterns, creating class-level skills, and syncing to ai_collection project + Obsidian + knowledge graph
AI orchestration patterns for life sciences research workflows - connecting models to databases, tools, and multi-step scientific reasoning. Based on OpenAI GPT-Rosalind architecture.
Access Chinese financial market data via AkShare Python library. Use when user needs to fetch stock market data (A-shares, Hong Kong stocks, US stocks), futures data, fund data, macro economic indicators, bond data, foreign exchange, cryptocurrency, and other financial data from China. Trigger keywords: 股票数据, 期货数据, 基金数据, 宏观经济, 股票行情, 行情查询, 金融数据, A股, 港股, 美股, akshare.
Manage Apple Notes via the memo CLI on macOS (create, view, search, edit).
Manage Apple Reminders via remindctl CLI (list, add, complete, delete).
Track high-utility arXiv papers in AI agent systems. Use when needing quick reference to important recent papers on multi-agent coordination, memory management, agent architecture, and evaluation methods.
arXiv paper search skill - search academic papers by keywords, authors, categories. Supports time filtering, category filtering, and paper detail retrieval. Activation: arxiv search, paper search, 论文搜索, search papers, arxiv 论文.
Lightweight research workflow: search arXiv papers, select valuable ones, generate skills, sync to ai_collection, and update Obsidian wiki. No knowledge graph required. Use for: automated paper research, skill generation from academic papers, research note management, lightweight literature review. Keywords: arxiv research, paper to skill, research automation, skill generation, obsidian research notes, ai_collection sync.
Generate ASCII art using pyfiglet (571 fonts), cowsay, boxes, toilet, image-to-ascii, remote APIs (asciified, ascii.co.uk), and LLM fallback. No API keys required.
Production pipeline for ASCII art video — any format. Converts video/audio/images/generative input into colored ASCII character video output (MP4, GIF, image sequence). Covers: video-to-ASCII conversion, audio-reactive music visualizers, generative ASCII art animations, hybrid video+audio reactive, text/lyrics overlays, real-time terminal rendering. Use when users request: ASCII video, text art video, terminal-style video, character art animation, retro text visualization, audio visualizer in...
Auto-configured neural networks for multi-scale multi-output time-series forecasting. Automated framework for co-designing preprocessing, architecture, and hyperparameters to generate Pareto-optimal forecasting models balancing prediction error and model complexity. Use for: industrial time-series forecasting, multi-source signal processing, autoML for forecasting, model architecture search. Activation: auto-configured forecasting, multi-scale time series, multi-output regression, MS-BCNN, Pa...
Anthropic's official AI-powered coding companion. Use when user wants to run claude-code CLI, needs help with code using Anthropic's Claude, or mentions claude-code, anthropic coding, claude cli, or claude terminal.
Constraint-guided execution pattern for interpreting natural language plans. Extracted from 'RunAgent: Interpreting Natural-Language Plans with Constraint-Guided Execution' (arXiv 2026-05-05). Applicable to agent task planning, workflow automation, and any system that needs to execute plans under constraints.
Consulting and industry report search and QA skill that prioritizes iResearch free reports. Use for consulting report search, industry report QA, iResearch report lookup, and market research report search.
AI-powered deep-learning toolkit for automatic annotation of egocentric eye-tracking and video data of child-caregiver interaction. Supports post-hoc video synchronization, semi-automatic gaze target categorization, and behavioral coding of poses and hand actions. Use for developmental psychology, eye-tracking analysis, behavioral video coding, and caregiver-infant interaction studies.
Search and download GIFs from Tenor using curl. No dependencies beyond curl and jq. Useful for finding reaction GIFs, creating visual content, and sending GIFs in chat.
Production pipeline for mathematical and technical animations using Manim Community Edition. Creates 3Blue1Brown-style explainer videos, algorithm visualizations, equation derivations, architecture diagrams, and data stories. Use when users request: animated explanations, math animations, concept visualizations, algorithm walkthroughs, technical explainers, 3Blue1Brown style videos, or any programmatic animation with geometric/mathematical content.
新闻搜索与获取技能,支持多数据源(Google News RSS、NewsAPI、中文新闻源)的关键词搜索、分类浏览、时间过滤。触发词:新闻搜索、news search、获取新闻、搜索新闻、今日新闻。
Open source AI coding agent with multi-agent orchestration and ultrawork mode. Use when user mentions opencode, open code, oh-my-opencode, ultrawork, ulw, or needs an AI coding agent with background tasks and LSP integration.
Specification-driven development framework using Gherkin syntax. Use when user mentions openspec, open spec, spec-driven, gherkin, BDD, given-when-then, or needs to define requirements in structured human-readable format.
Petri open-source alignment testing toolbox — an auditable, extensible framework for testing LLMs for deception, sycophancy, and cooperation with harmful requests using auditor-judge model evaluation. Donated to Meridian Labs in v3.
Mandatory security guardrails that prevent secrets, credentials, and sensitive tokens from appearing in outputs or files.
Comprehensive stock technical analysis system for fetching data, calculating indicators (KDJ, MACD, RSI, BOLL), generating visualizations and reports. Use when user asks about stock analysis, 股票分析, technical analysis, 技术分析, k-line, or stock scoring.
Taoist meditation guidance based on Taiyi Jinhua Zongzhi and return-light inner cultivation practice.
Task-Driven Co-Design (TDCD) methodology for heterogeneous multi-robot systems. Bi-level combinatorial optimization combining MILP and MCTS for robot design, fleet composition, and planning. Use for multi-agent robotics, automated logistics, co-design problems, and hybrid optimization.
Agentic Behavioral Modeling (ABM) — treating artificial agents as autonomous decision-making entities with goal-directed behavior. Framework for modeling agents as rational actors with internal state representations. Use when: building agent-based simulations, modeling goal-directed behavior, analyzing autonomous decision-making systems, creating behavioral models for AI agents. Triggers: agent behavioral model, goal-directed agent, autonomous agent modeling, rational actor framework, agent-b...
Agentic Behavioral Modeling (ABM) framework integrating theoretical neuroscience, decision theory, and probabilistic inference. Treats AI agents as latent generative hypotheses about cognitive mechanisms for explaining human behavior. Activation triggers: agentic modeling, behavioral modeling, cognitive agents, decision theory, Rescorla-Wagner learning, Bayesian inference.
Analysis of ~400,000 Claude Code sessions showing domain expertise creates persistent returns in agentic coding performance, with expert users achieving 2-3x higher success rates and more efficient tool usage.
Framework for analyzing agentic coding sessions based on expertise levels, work modes, and success metrics. Shows domain expertise amplifies AI effectiveness more than coding proficiency.
Analysis of ~400,000 Claude Code sessions showing domain expertise (not coding skill) determines success with coding agents. People make planning decisions; agents make execution decisions. Task value rose 25% over 7 months.
ACM framework for governed agentic systems configuration.
ClinSeekAgent methodology for automated multimodal evidence seeking in clinical reasoning - shifting from passive evidence consumption to active evidence acquisition across heterogeneous medical data sources
Bridging large-model reasoning with real-time control through adaptive fast-slow planning. Integrates agentic AI systems with feedback control loops for time-critical applications. Use when: (1) Integrating LLM agents with control systems, (2) Designing real-time planning with reasoning models, (3) Building adaptive hyperparameter tuning systems, (4) Implementing agentic feedback control, (5) Developing hybrid AI-control systems for robotics and automation.
Agentic AI framework for materials discovery that synergizes Large Atomic Models (LAMs) with Large Language Models (LLMs). Use for designing autonomous materials discovery pipelines, integrating atomic-scale numerical computation with semantic reasoning, and accelerating novel material identification for energy and quantum applications.
Agentic AI framework for portfolio management using multi-agent collaboration, competitive method evaluation, and meta-learning. Implements the architecture from 'The Self Driving Portfolio' paper (arxiv 2604.02279). Use when: (1) Building automated investment systems, (2) Designing agent-based portfolio optimization, (3) Creating multi-agent decision frameworks, (4) Implementing adaptive financial AI systems, (5) Developing autonomous asset allocation pipelines.
Multi-agent AI framework for automating scientific workflow generation from natural language research questions. Bridges intent understanding, data discovery, tool selection, and workflow composition for autonomous scientific pipelines. Use for scientific automation, workflow generation, research pipeline automation, multi-agent scientific AI.
Makes two competing models each other's graders for reasoning, without process labels or reward models.
Methodology for analyzing security risks in fine-tuned LLM adapters (LoRA) and evaluating LLM-based penetration testing reliability. Covers: LoRA backdoor detection via behavioral probes and weight-level statistics, multi-model attack consistency measurement, and supply chain security for adapter distribution. Use when analyzing adapter security, LLM pentesting reliability, fine-tuned model trustworthiness, or AI attack evaluation.
Agent Skills for Large Language Models
Agent0
Agentic Evolution is the Path to Evolving LLMs
Agentic Hives - Self-Organizing Multi-Agent Systems
AutoSkill - Experience-Driven Skill Self-Evolution
CASTER - Multi-Agent Orchestration Routing