All authors

Claude Skills by hiyenwong
github.com/hiyenwong9,934 skills5 installs19,223 views
- React Revealing Evolutionary Action Consequence Trajectories For Interpretable Reinforcement Learning**arXiv ID:** 2404.03359 **Authors:** Philipp Altmann, Céline Davignon, Maximilian Zorn, Fabian Ritz, Claudia Linnhoff-Popien, Thomas Gabor **Published:** 2024-04-04T10:56:30Z **Abstract:** To enhance the interpretability of Reinforcement Learning (RL), we propose Revealing Evolutionary Action Consequence Trajectories (REACT). In contrast to the prevalent practice of validating RL models based on their optimal behavior learned during training, we posit that considering a range of edge-case tr...Votes: 0GitHub stars: 3
- Real World Evaluation Ai Agent Drafting Translational Impact SummariesSkill derived from arXiv:2607.16989 - Real-World Evaluation of an AI Agent Drafting Translational Impact SummariesVotes: 0GitHub stars: 3
- Real World Evaluation Of An Ai Agent Drafting TranDerived from arXiv:2607.16989 - Real-World Evaluation of an AI Agent Drafting Translational Impact SummariesVotes: 0GitHub stars: 3
- Realworld Validation Of Safe Reinforcement Learning Model Predictive Control And Decision Treebased Home Energy Management Systems**arXiv ID:** 2408.07435 **Authors:** Julian Ruddick, Glenn Ceusters, Gilles Van Kriekinge, Evgenii Genov, Cedric De Cauwer, Thierry Coosemans, Maarten Messagie **Published:** 2024-08-14T10:12:15Z **Abstract:** Recent advancements in machine learning based energy management approaches, specifically reinforcement learning with a safety layer (OptLayerPolicy) and a metaheuristic algorithm generating a decision tree control policy (TreeC), have shown promise. However, their effectiveness has onl...Votes: 0GitHub stars: 3
- Recursive Self Improvement In Ai From Bounded Self Refinement To AutonomousAI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data they generate, and, increasingly, conducting AI r. Based on arXiv:2607.07663.Votes: 0GitHub stars: 3
- Reducing Catastrophic Forgetting When Evolving Neural Networks**arXiv ID:** 1904.03178 **Authors:** Joseph Early **Published:** 2019-04-05T17:57:29Z **Abstract:** A key stepping stone in the development of an artificial general intelligence (a machine that can perform any task), is the production of agents that can perform multiple tasks at once instead of just one. Unfortunately, canonical methods are very prone to catastrophic forgetting (CF) - the act of overwriting previous knowledge about a task when learning a new task. Recent efforts have develop...Votes: 0GitHub stars: 3
- Regulating Autonomous And Agentic AiRegulating autonomous and agentic AIVotes: 0GitHub stars: 3
- Reinforcement Learning Algorithms Foundation ModelsSkill derived from arXiv:2607.17560 - Reinforcement Learning: From Algorithms To Foundation ModelsVotes: 0GitHub stars: 3
- Reinforcement Learning From Algorithms To FoundatiDerived from arXiv:2607.17560 - Reinforcement Learning: From Algorithms To Foundation ModelsVotes: 0GitHub stars: 3
- Reinforcement Learning With Chromatic Networks For Compact Architecture Search**arXiv ID:** 1907.06511 **Authors:** Xingyou Song, Krzysztof Choromanski, Jack Parker-Holder, Yunhao Tang, Wenbo Gao, Aldo Pacchiano, Tamas Sarlos, Deepali Jain, Yuxiang Yang **Published:** 2019-07-10T16:57:50Z **Abstract:** We present a neural architecture search algorithm to construct compact reinforcement learning (RL) policies, by combining ENAS and ES in a highly scalable and intuitive way. By defining the combinatorial search space of NAS to be the set of different edge-partitionings (...Votes: 0GitHub stars: 3
- Reinforcement Learning With Prediction Based RewarSkill for AI agent capabilitiesVotes: 0GitHub stars: 3
- Distributed Zeroth Order MarlhDistributed zeroth-order policy gradient for networked multi-agent reinforcement learning from human feedback (RLHF). Addresses scalability of preference-based RL to multi-agent systems without centralized training. Activation: distributed multi-agent RLHF, zeroth-order policy gradient, networked MARL, human feedback multi-agent, decentralized preference RL, spatial-temporal truncated trajectory.Votes: 0GitHub stars: 3
- Ohp Rl Human Preference GuidanceOHP-RL methodology — using online human preference interventions to guide reinforcement learning policy for robot manipulation. Addresses unsafe exploration in real-world RL by encoding human interventions as relative preference signals. Activation: OHP-RL, online human preference, human-in-the-loop RL, robot manipulation RL, human-guided RL, human intervention RL.Votes: 0GitHub stars: 3
- Retain Consolidate Budget Dependent Operator Selection Language Agent MemorySkill derived from arXiv:2607.17545 - Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent MemoryVotes: 0GitHub stars: 3
- Reusability And Transferability Of Macro Actions For Reinforcement Learning**arXiv ID:** 1908.01478 **Authors:** Yi-Hsiang Chang, Kuan-Yu Chang, Henry Kuo, Chun-Yi Lee **Published:** 2019-08-05T05:59:40Z **Abstract:** Conventional reinforcement learning (RL) typically determines an appropriate primitive action at each timestep. However, by using a proper macro action, defined as a sequence of primitive actions, an agent is able to bypass intermediate states to a farther state and facilitate its learning procedure. The problem we would like to investigate is what ass...Votes: 0GitHub stars: 3
- Reward Learning From Human Preferences And Demonstrations In Atari**arXiv ID:** 1811.06521 **Authors:** Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, Dario Amodei **Published:** 2018-11-15T18:33:43Z **Abstract:** To solve complex real-world problems with reinforcement learning, we cannot rely on manually specified reward functions. Instead, we can have humans communicate an objective to the agent directly. In this work, we combine two approaches to learning from human feedback: expert demonstrations and trajectory preferences. We train...Votes: 0GitHub stars: 3
- Rl Compositional Reasoning StrategiesUnderstanding and leveraging how RL composes primitive skills into higher-level reasoning strategies.Votes: 0GitHub stars: 3
- Rl Ion ShuttlingReinforcement learning for ion shuttling optimization on trapped-ion quantum computersVotes: 0GitHub stars: 3
- Rl Neural Model EditingReinforcement learning framework for neural model editing where agents learn to modify models via reward feedback instead of manually engineered algorithmsVotes: 0GitHub stars: 3
- Rl Nqs OptimizationFrame neural quantum state optimization as reinforcement learning for scalable wavefunction approximation.Votes: 0GitHub stars: 3
- Rl Temporal LogicCombine reinforcement learning with signal temporal logic (STL) for stratified control. Use STL specifications to define complex temporal constraints and stratification for hierarchical RL policy learning. Activation: RL temporal logic, STL reinforcement learning, temporal specification RL, stratified control.Votes: 0GitHub stars: 3
- Rl Triton Gpu Kernels Credit AssignmentHigh-performance RL credit assignment via Triton kernels.Votes: 0GitHub stars: 3
- Rl Tsch Dynamic ListeningReinforcement Learning-driven Adaptive Listening for TSCH NetworksVotes: 0GitHub stars: 3
- Rl² Fast Reinforcement Learning Via Slow ReinforceSkill for AI agent capabilitiesVotes: 0GitHub stars: 3
- Robot Co Design Inductive BiasesInductive biases identification for morphology-control co-design in robotics. Analyzes co-design landscapes to discover low-dimensional manifolds and patterns for sample-efficient search. Activation: robot co-design, morphology optimization, control co-design, inductive biases, high-dimensional search, soft robotics, embodied AI.Votes: 0GitHub stars: 3
- Roser Rl Component SynergyROSER for RL component synergy in sample-efficient control.Votes: 0GitHub stars: 3
- Rp Regret Adaptive OpponentsRepeated Policy Regret (RP-Regret) methodology for regret minimization in repeated games with adaptive opponents — addresses limitations of external regret when opponents respond to history of play.Votes: 0GitHub stars: 3
- Rt Shcua Real Time Self Hosted Computer Use Agent Uav ControlSkill derived from arXiv:2607.17951 - RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV ControlVotes: 0GitHub stars: 3
- Saas Self Aware Agentic SearchSelf-Aware Reinforcement Learning for Over-Search Mitigation - dynamic self-awareness that regulates search behavior without compromising accuracyVotes: 0GitHub stars: 3
- Safe Rl Forward InvariantLearning over Forward-Invariant Policy Classes: Reinforcement Learning without Safety Concerns. Novel safe RL framework embedding safety directly into action representation via forward-invariance-induced action-space design. Finite admissible actions correspond to stabilizing feedback laws preserving forward invariance of safe state set. Decouples safety assurance from performance optimization. Use for: (1) Safe reinforcement learning, (2) forward-invariant set design, (3) safety-by-construct...Votes: 0GitHub stars: 3
- Saga Synthetic Agentic Graph Architecture For TempDerived from arXiv:2607.17288 - SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark GenerationVotes: 0GitHub stars: 3
- Saga Synthetic Agentic Graph Architecture Temporal Benchmark GenerationSkill derived from arXiv:2607.17288 - SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark GenerationVotes: 0GitHub stars: 3
- Sao Single Rollout Async RlSingle-rollout Asynchronous Optimization (SAO) methodology for agentic RL. Replaces GRPO's group-wise sampling with single-rollout sampling to reduce off-policy effects and improve stability in asynchronous training. Successfully deployed for GLM-5.2 (750B-A40B).Votes: 0GitHub stars: 3
- Scalelogic Rl ReasoningMethodology for studying RL scaling laws in LLM reasoning using a synthetic logical reasoning framework (ScaleLogic) with independent control over proof depth and logical expressiveness.Votes: 0GitHub stars: 3
- Scaling Laws For Reward Model OveroptimizationSkill for AI agent capabilitiesVotes: 0GitHub stars: 3
- Scaling Self Play With Self GuidanceSelf-Guided Self-Play (SGS) prevents conjecturer collapse in LLM self-play.Votes: 0GitHub stars: 3
- Sd Search On Policy Hindsight DistillationSD-Search methodology for search-augmented reasoning. Derives step-level supervision from the policy itself through on-policy hindsight self-distillation, without external teacher or annotations.Votes: 0GitHub stars: 3
- Sdar Self Distilled Agentic RlSelf-Distilled Agentic Reinforcement Learning (SDAR) methodology. Stabilizes on-policy self-distillation (OPSD) for multi-turn LLM agents by treating distillation as a gated auxiliary while keeping RL as the primary backbone. Maps detached token-level signals into a sigmoid gate, strengthening distillation on positive-gap tokens and attenuating negative teacher rejections. Use when designing RL post-training for LLM agents, combining OPSD with GRPO/PPO, or addressing multi-turn distillation i...Votes: 0GitHub stars: 3
- Self Evolving Agent ExperienceFramework for LLM agents that accumulates and reuses cross-task experience including verified skills, statistical evidence of effective strategies, and recurring error-fix patterns. Enables zero-test-time search on new tasks.Votes: 0GitHub stars: 3
- Self Evolving Agents SurveySkill for AI agent capabilitiesVotes: 0GitHub stars: 3
- Self Organizing Classifiers First Steps In Structured Evolutionary Machine Learning**arXiv ID:** 1811.08225 **Authors:** Danilo Vasconcellos Vargas, Hirotaka Takano, Junichi Murata **Published:** 2018-11-20T13:00:51Z **Abstract:** Learning classifier systems (LCSs) are evolutionary machine learning algorithms, flexible enough to be applied to reinforcement, supervised and unsupervised learning problems with good performance. Recently, self organizing classifiers were proposed which are similar to LCSs but have the advantage that in its structured population no balance betwe...Votes: 0GitHub stars: 3
- Selfevidencing Through Hierarchical Gradient Decomposition A Dissipative System That Maintains Nonequilibrium Steadystate By Minimizing Variational Free Energy**arXiv ID:** 2510.17916 **Authors:** Michael James McCulloch **Published:** 2025-10-20T00:19:32Z **Abstract:** The Free Energy Principle (FEP) states that self-organizing systems must minimize variational free energy to persist, but the path from principle to implementable algorithm has remained unclear. We present a constructive proof that the FEP can be realized through exact local credit assignment. The system decomposes gradient computation hierarchically: spatial credit via feedback ali...Votes: 0GitHub stars: 3
- Semiconductor Fab Scheduling With Selfsupervised And Reinforcement Learning**arXiv ID:** 2302.07162 **Authors:** Pierre Tassel, Benjamin Kovács, Martin Gebser, Konstantin Schekotihin, Patrick Stöckermann, Georg Seidel **Published:** 2023-02-14T16:15:50Z **Abstract:** Semiconductor manufacturing is a notoriously complex and costly multi-step process involving a long sequence of operations on expensive and quantity-limited equipment. Recent chip shortages and their impacts have highlighted the importance of semiconductors in the global supply chains and how reliant on...Votes: 0GitHub stars: 3
- Shortcut Trajectory Planning Offline RlSingle-stage shortcut-model trajectory planner for efficient offline model-based RL. Replaces two-stage consistency-distillation with one-stage step-size-conditioned shortcut models + feasibility-aware critic for fast one/few-step generative planning. Use when building fast diffusion-based planners for offline RL (D4RL) without the training cost/instability of teacher-student distillation.Votes: 0GitHub stars: 3
- Silent Failures In Multimodal Agentic Searcha DiagSkill generated from arXiv paper 2607.19793: Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge EvaluationVotes: 0GitHub stars: 3
- Sim To Real Transfer Of Robotic Control With DynamSkill for AI agent capabilitiesVotes: 0GitHub stars: 3
- Simultaneously Evolving Deep Reinforcement Learning Models Using Multifactorial Optimization**arXiv ID:** 2002.12133 **Authors:** Aritz D. Martinez, Eneko Osaba, Javier Del Ser, Francisco Herrera **Published:** 2020-02-25T10:36:57Z **Abstract:** In recent years, Multifactorial Optimization (MFO) has gained a notable momentum in the research community. MFO is known for its inherent capability to efficiently address multiple optimization tasks at the same time, while transferring information among such tasks to improve their convergence speed. On the other hand, the quantum leap made ...Votes: 0GitHub stars: 3
- Single Rollout Asynchronous Optimization For Agentic Reinforcement LearningReinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is in. Based on arXiv:2607.07508.Votes: 0GitHub stars: 3
- Skillcenter A Large Scale Source Grounded Skill Library For Autonomous Ai AgentsAutonomous AI agents can execute complex tasks with limited human review, yet they often lack the grounded operational knowledge to make their outputs not just executable but correct, secure, and main. Based on arXiv:2607.07676.Votes: 0GitHub stars: 3
- Skillrise Agentic Reinforcement Learning For CrossSkillRise: Agentic Reinforcement Learning for Cross-Task Skill EvolutionVotes: 0GitHub stars: 3