Derived from arXiv:2607.17760 - Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning
Scanned 9/11/2026
Install to Claude Code
npx -y skills add hiyenwong/ai_collection --skill generalize-and-guide-decomposing-rewards-for-few-s --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Generalize And Guide Decomposing Rewards For Few S?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/hiyenwong-generalize-and-guide-decomposing-rewards-for-few-s)More formats (shields.io, HTML) on the badges page.
# Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning
Derived from arXiv:2607.17760 - Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning
## Core Concept
Inverse reinforcement learning (IRL) provides a powerful framework for learning from demonstrations. However, real-world tasks often exhibit substantial natural variations (e.g., picking up mugs with varying shapes), making it impractical to collect demonstrations that fully specify a new task under every possible scenario. In practice, while demonstrations for the target task are limited, it is often easier to obtain datasets of heterogeneous but related behaviors. This motivates the problem of...
## Key Insights
- Derived from arXiv:2607.17760
- Published: 2026-07-20
- Utility Score: 1.00
- Authors: Ziyi Liu, Grace Zhang
## Activation
generalize-and-guide-decomposing-rewards-for-few-s, 2607.17760
## References
- arXiv: https://arxiv.org/abs/2607.17760
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!
Understand the components, mechanics, and constraints of context in agent systems. Use when designing agent architectures, debugging context-related failures, or optimizing context usage.
Draft release notes and changelog entries from git history or merged PRs between two refs (tags/SHAs/branches), including breaking changes, migrations, and upgrade steps. Use when the user asks for release notes, changelog updates, or a GitHub Release draft.
Documentation style guide enforcer by @planetabhi. Applies and reviews the writing style guide when authoring or editing product documentation and tutorials. Use to check prose for voice, tense, word choice, inclusive language, formatting, code block, UI, Markdown, and number/date conventions.
Quick-reference card for all caveman modes, skills, and commands. One-shot display, not a persistent mode. Trigger: /caveman-help, "caveman help", "what caveman commands", "how do I use caveman".
Explain how claude-mem captures observations, when memory injection kicks in, and where data lives. Use when the user asks "how does claude-mem work?" or "what is this thing doing?".