All authors

Claude Skills by hiyenwong
github.com/hiyenwong9,934 skills5 installs19,223 views
- Socialgrid Embodied Multi AgentResearch paper: SocialGrid - A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems. Among Us-inspired environment for social reasoning evaluation.Votes: 0GitHub stars: 3
- Soft Actorcritic With Inhibitory Networks For Faster Retraining**arXiv ID:** 2202.02918 **Authors:** Jaime S. Ide, Daria Mićović, Michael J. Guarino, Kevin Alcedo, David Rosenbluth, Adrian P. Pope **Published:** 2022-02-07T03:10:34Z **Abstract:** Reusing previously trained models is critical in deep reinforcement learning to speed up training of new agents. However, it is unclear how to acquire new skills when objectives and constraints are in conflict with previously learned skills. Moreover, when retraining, there is an intrinsic conflict between explo...Votes: 0GitHub stars: 3
- Soft Control Multi AgentSoft Control methodology for guiding collective behavior in multi-agent systems using shill agents. Based on Han et al. (2010) research on controlling self-organized systems without modifying local agent rules.Votes: 0GitHub stars: 3
- Solving Rubiks Cube With A Robot HandSkill for AI agent capabilitiesVotes: 0GitHub stars: 3
- Spacellagent A Self Evolving Llm Based Multi Agent Framework For TrajectorySpatial and Single-cell transcriptomics are transformative in deciphering cellular dynamics. As the fundamental paradigm for reconstructing cell developmental paths, trajectory inference (TI) is criti. Based on arXiv:2607.07467.Votes: 0GitHub stars: 3
- Spacemind A Modular And Self Evolving Embodied VisResearch paper: SpaceMind: A Modular and Self-Evolving Embodied Vision-Language Agent FrameworkVotes: 0GitHub stars: 3
- Sparse Evidence Can Suffice Agentic Evidence SeekiDerived from arXiv:2607.18080 - Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation DetectionVotes: 0GitHub stars: 3
- Sparse Evidence Suffice Agentic Evidence Seeking Multimodal Video Misinformation DetectionSkill derived from arXiv:2607.18080 - Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation DetectionVotes: 0GitHub stars: 3
- Sparse Meta Networks For Sequential Adaptation And Its Application To Adaptive Language Modelling**arXiv ID:** 2009.01803 **Authors:** Tsendsuren Munkhdalai **Published:** 2020-09-03T17:06:52Z **Abstract:** Training a deep neural network requires a large amount of single-task data and involves a long time-consuming optimization phase. This is not scalable to complex, realistic environments with new unexpected changes. Humans can perform fast incremental learning on the fly and memory systems in the brain play a critical role. We introduce Sparse Meta Networks -- a meta-learning approach ...Votes: 0GitHub stars: 3
- Sparse Training Theory For Scalable And Efficient Agents**arXiv ID:** 2103.01636 **Authors:** Decebal Constantin Mocanu, Elena Mocanu, Tiago Pinto, Selima Curci, Phuong H. Nguyen, Madeleine Gibescu, Damien Ernst, Zita A. Vale **Published:** 2021-03-02T10:48:29Z **Abstract:** A fundamental task for artificial intelligence is learning. Deep Neural Networks have proven to cope perfectly with all learning paradigms, i.e. supervised, unsupervised, and reinforcement learning. Nevertheless, traditional deep learning approaches make use of cloud computing...Votes: 0GitHub stars: 3
- Specifying Delegated Autonomy Boundary Requirements Engineering Agentic AiSkill derived from arXiv:2607.17225 - Specifying the Delegated-Autonomy Boundary: Requirements Engineering for Agentic AIVotes: 0GitHub stars: 3
- Sr Agent An Experience Driven Agentic Framework FoDerived from arXiv:2607.17719 - SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategies Refinement in E-Commerce RecommendationVotes: 0GitHub stars: 3
- Sr Agent Experience Driven Agentic Framework Post Ranking Strategies Refinement E CommercSkill derived from arXiv:2607.17719 - SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategies Refinement in E-CommercVotes: 0GitHub stars: 3
- Statefuse Conflict Preserving Memory Multi AgentConflict-aware replicated memory contract for multi-agent systems. Agent systems accumulate conflicting observations across branches, retries, and replicas. StateFuse builds a conflict-preserving memory layer on standard version control principles, avoiding information loss from overwrite rules. Activation: StateFuse, conflict-aware memory, replicated memory, multi-agent memory, conflict-preserving, version control memory, agent memory branches.Votes: 0GitHub stars: 3
- Statistical Efficiency Quantile Distributional Reinforcement LearningStudies quantile-based distributional RL from statistical efficiency perspective. Non-asymptotic error bound O(√(m/n)) under W∞ metric. Achieves optimal √n convergence rate. Asymptotic distribution and semiparametric efficiency bound. Berry-Esseen theorem. Activation: distributional RL, quantile regression, statistical efficiency, policy evaluation, return distribution.Votes: 0GitHub stars: 3
- Steppo Agentic RlStepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning - A novel RL framework for training LLM agents with step-level credit assignmentVotes: 0GitHub stars: 3
- Stp Stabilizes Goal Conditioned DynamicsShort-Term Synaptic Plasticity (STP) stabilizes goal-conditioned dynamics in PFC-inspired reservoir computing models for multistep goal-directed action planning. Combines STP with basal-ganglia-inspired temporal-difference learning. Achieves 89.2% success under noise (vs 49.5% without STP). Activation: synaptic plasticity, reservoir computing, goal-conditioned dynamics, PFC model, action planning, goal-directed behavior, temporal-difference learning, effective connectivity, short-term plastic...Votes: 0GitHub stars: 3
- Switchbased Active Deep Dynaq Efficient Adaptive Planning For Taskcompletion Dialogue Policy Learning**arXiv ID:** 1811.07550 **Authors:** Yuexin Wu, Xiujun Li, Jingjing Liu, Jianfeng Gao, Yiming Yang **Published:** 2018-11-19T08:23:34Z **Abstract:** Training task-completion dialogue agents with reinforcement learning usually requires a large number of real user experiences. The Dyna-Q algorithm extends Q-learning by integrating a world model, and thus can effectively boost training efficiency using simulated experiences generated by the world model. The effectiveness of Dyna-Q, however, dep...Votes: 0GitHub stars: 3
- Synergizing Qualitydiversity With Descriptorconditioned Reinforcement Learning**arXiv ID:** 2401.08632 **Authors:** Maxence Faldor, Félix Chalumeau, Manon Flageat, Antoine Cully **Published:** 2023-12-10T19:53:15Z **Abstract:** A hallmark of intelligence is the ability to exhibit a wide range of effective behaviors. Inspired by this principle, Quality-Diversity algorithms, such as MAP-Elites, are evolutionary methods designed to generate a set of diverse and high-fitness solutions. However, as a genetic algorithm, MAP-Elites relies on random mutations, which can become...Votes: 0GitHub stars: 3
- Terminal Curl Arxiv Security BlockTerminal curl to arXiv API blocked by security scanner (exit -1) — new failure mode 2026-06-02. Different from plain HTTP or pipe-to-interpreter blocks. Workaround: browser_navigate. Activation: terminal curl blocked, arxiv security block, curl exit -1, arxiv api blocked.Votes: 0GitHub stars: 3
- Terrazero Procedural Driving Simulation For Zero DTerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale - Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground beh...Votes: 0GitHub stars: 3
- The Blind Curator How A Biased Judge Silently Disables Skill Retirement In SelfA self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the structural constraint that keeps a growing library f. Based on arXiv:2607.07436.Votes: 0GitHub stars: 3
- The Ethics Of Autonomous Ai Agents For Offensive SSkill generated from arXiv paper 2607.20255: The Ethics of Autonomous AI Agents for Offensive SecurityVotes: 0GitHub stars: 3
- Think Big Search Small Where Capacity Matters In Hierarchical Search AgentsLarge language model based search agents increasingly adopt multi-agent architectures in which a main agent decomposes a complex question into sub-queries and dispatches them to parallel sub-agents. H. Based on arXiv:2607.07548.Votes: 0GitHub stars: 3
- Third Person Imitation LearningSkill for AI agent capabilitiesVotes: 0GitHub stars: 3
- Total Stochastic Gradient Algorithms And Applications In Reinforcement Learning**arXiv ID:** 1902.01722 **Authors:** Paavo Parmas **Published:** 2019-02-05T14:54:05Z **Abstract:** Backpropagation and the chain rule of derivatives have been prominent; however, the total derivative rule has not enjoyed the same amount of attention. In this work we show how the total derivative rule leads to an intuitive visual framework for creating gradient estimators on graphical models. In particular, previous "policy gradient theorems" are easily derived. We derive new gradient estima...Votes: 0GitHub stars: 3
- Toward Preferencealigned Large Language Models Via Residualbased Model Steering**arXiv ID:** 2509.23982 **Authors:** Lucio La Cava, Andrea Tagarelli **Published:** 2025-09-28T17:16:16Z **Abstract:** Preference alignment is a critical step in making Large Language Models (LLMs) useful and aligned with (human) preferences. Existing approaches such as Reinforcement Learning from Human Feedback or Direct Preference Optimization typically require curated data and expensive optimization over billions of parameters, and eventually lead to persistent task-specific models. In th...Votes: 0GitHub stars: 3
- Towards Agentic Agent Based Models Feasibility PerDerived from arXiv:2607.17948 - Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model CheckingVotes: 0GitHub stars: 3
- Towards Agentic Agent Based Models Feasibility Performance Statistical Model CheckingSkill derived from arXiv:2607.17948 - Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model CheckingVotes: 0GitHub stars: 3
- Towards Empathic Deep Qlearning**arXiv ID:** 1906.10918 **Authors:** Bart Bussmann, Jacqueline Heinerman, Joel Lehman **Published:** 2019-06-26T08:59:02Z **Abstract:** As reinforcement learning (RL) scales to solve increasingly complex tasks, interest continues to grow in the fields of AI safety and machine ethics. As a contribution to these fields, this paper introduces an extension to Deep Q-Networks (DQNs), called Empathic DQN, that is loosely inspired both by empathy and the golden rule ("Do unto others as you would ha...Votes: 0GitHub stars: 3
- Towards Humanlike Driving Active Inference In Autonomous Vehicle Control**arXiv ID:** 2407.07684 **Authors:** Elahe Delavari, John Moore, Junho Hong, Jaerock Kwon **Published:** 2024-07-10T14:08:27Z **Abstract:** This paper presents a novel approach to Autonomous Vehicle (AV) control through the application of active inference, a theory derived from neuroscience that conceptualizes the brain as a predictive machine. Traditional autonomous driving systems rely heavily on Modular Pipelines, Imitation Learning, or Reinforcement Learning, each with inherent limitatio...Votes: 0GitHub stars: 3
- Towards Mental Time Travel A Hierarchical Memory For Reinforcement Learning Agents**arXiv ID:** 2105.14039 **Authors:** Andrew Kyle Lampinen, Stephanie C. Y. Chan, Andrea Banino, Felix Hill **Published:** 2021-05-28T18:12:28Z **Abstract:** Reinforcement learning agents often forget details of the past, especially after delays or distractor tasks. Agents with common memory architectures struggle to recall and integrate across multiple timesteps of a past event, or even to recall the details of a single timestep that is followed by distractor tasks. To address these limitati...Votes: 0GitHub stars: 3
- Towards Robust And Domain Agnostic Reinforcement Learning Competitions**arXiv ID:** 2106.03748 **Authors:** William Hebgen Guss, Stephanie Milani, Nicholay Topin, Brandon Houghton, Sharada Mohanty, Andrew Melnik, Augustin Harter, Benoit Buschmaas, Bjarne Jaster, Christoph Berganski, Dennis Heitkamp, Marko Henning, Helge Ritter, Chengjie Wu, Xiaotian Hao, Yiming Lu, Hangyu Mao, Yihuan Mao, Chao Wang, Michal Opanowicz, Anssi Kanervisto, Yanick Schraner, Christian Scheller, Xiren Zhou, Lu Liu, Daichi Nishio, Toi Tsuneda, Karolis Ramanauskas, Gabija Juceviciute **P...Votes: 0GitHub stars: 3
- Towards Self Evolving Agents A Human Inspired AdapSkill derived from arXiv paper 2607.11913: Towards Self-Evolving Agents: A Human-Inspired Adaptive Exploration-Exploitation Framework for Genetic Network ProgrammingVotes: 0GitHub stars: 3
- Training Data Set Assessment For Decisionmaking In A Multiagent Landmine Detection Platform**arXiv ID:** 2004.05380 **Authors:** Johana Florez-Lozano, Fabio Caraffini, Carlos Parra, Mario Gongora **Published:** 2020-04-11T12:05:30Z **Abstract:** Real-world problems such as landmine detection require multiple sources of information to reduce the uncertainty of decision-making. A novel approach to solve these problems includes distributed systems, as presented in this work based on hardware and software multi-agent systems. To achieve a high rate of landmine detection, we evaluate th...Votes: 0GitHub stars: 3
- Trajonco Multi Agent Framework Temporal ReasoningAccurate estimation of cancer risk from longitudinal electronic health records (EHRs) could support earlier detection and improved care, but modeling ... 触发词: 多智能体系统, 控制系统.Votes: 0GitHub stars: 3
- Treeqn And Atreec Differentiable Treestructured Models For Deep Reinforcement Learning**arXiv ID:** 1710.11417 **Authors:** Gregory Farquhar, Tim Rocktäschel, Maximilian Igl, Shimon Whiteson **Published:** 2017-10-31T11:54:35Z **Abstract:** Combining deep model-free reinforcement learning with on-line planning is a promising approach to building on the successes of deep RL. On-line planning with look-ahead trees has proven successful in environments where transition models are known a priori. However, in complex environments where transition models need to be learned from data...Votes: 0GitHub stars: 3
- Trim Reducing Ai Generated Codeslop Agent Trajectory MinimizationSkill derived from arXiv:2607.18161 - TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory MinimizationVotes: 0GitHub stars: 3
- Trim Reducing Ai Generated Codeslop Via Agent TrajDerived from arXiv:2607.18161 - TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory MinimizationVotes: 0GitHub stars: 3
- Trustworthy Ai For Process Automation On A Chyllahaase Polymerization Reactor**arXiv ID:** 2108.13381 **Authors:** Daniel Hein, Daniel Labisch **Published:** 2021-08-30T17:04:04Z **Abstract:** In this paper, genetic programming reinforcement learning (GPRL) is utilized to generate human-interpretable control policies for a Chylla-Haase polymerization reactor. Such continuously stirred tank reactors (CSTRs) with jacket cooling are widely used in the chemical industry, in the production of fine chemicals, pigments, polymers, and medical products. Despite appearing rathe...Votes: 0GitHub stars: 3
- Two Calls Beat Five Agents Evaluating Multi AgentTwo Calls Beat Five Agents: Evaluating Multi-Agent Pipelines Against Self-Refinement for Local LanguVotes: 0GitHub stars: 3
- Unbiased Recovery Policy GradientSafeExplorer - drop-in PPO modification for RL with a deterministic recovery policy that prevents silent bias in on-policy updates. Uses score-function estimator only at safe timesteps, never evaluates recovery-policy density, so it stays valid where importance sampling breaks. Use when training RL agents on real robots / safety-critical systems where falls/crashes are costly and a separate recovery controller is engaged inside the safe region.Votes: 0GitHub stars: 3
- Unfireable Safety Kernel Agent AlignmentExecution-time AI alignment architecture using an unfireable safety kernel that operates outside the agent's address space. Ensures safety controls cannot be bypassed by the AI agent itself, addressing the fundamental vulnerability of in-process guardrails.Votes: 0GitHub stars: 3
- Unsupervised Learning And Exploration Of Reachable Outcome Space**arXiv ID:** 1909.05508 **Authors:** Giuseppe Paolo, Alban Laflaquière, Alexandre Coninx, Stephane Doncieux **Published:** 2019-09-12T08:47:44Z **Abstract:** Performing Reinforcement Learning in sparse rewards settings, with very little prior knowledge, is a challenging problem since there is no signal to properly guide the learning process. In such situations, a good search strategy is fundamental. At the same time, not having to adapt the algorithm to every single problem is very desirable...Votes: 0GitHub stars: 3
- Unsupervised Learning Of Temporal Abstractions With Slotbased Transformers**arXiv ID:** 2203.13573 **Authors:** Anand Gopalakrishnan, Kazuki Irie, Jürgen Schmidhuber, Sjoerd van Steenkiste **Published:** 2022-03-25T10:59:46Z **Abstract:** The discovery of reusable sub-routines simplifies decision-making and planning in complex reinforcement learning problems. Previous approaches propose to learn such temporal abstractions in a purely unsupervised fashion through observing state-action trajectories gathered from executing a policy. However, a current limitation is t...Votes: 0GitHub stars: 3
- Unveiling Complex Collective Behaviors From SimpleUnveiling Complex Collective Behaviors from Simple Rewards - Multi-agent Reinforcement Learning (MARL) holds great potential for robot swarms, but the black-box nature of neural policies complicates strategic an...Votes: 0GitHub stars: 3
- Unveiling The Decisionmaking Process In Reinforcement Learning With Genetic Programming**arXiv ID:** 2407.14714 **Authors:** Manuel Eberhardinger, Florian Rupp, Johannes Maucher, Setareh Maghsudi **Published:** 2024-07-20T00:45:03Z **Abstract:** Despite tremendous progress, machine learning and deep learning still suffer from incomprehensible predictions. Incomprehensibility, however, is not an option for the use of (deep) reinforcement learning in the real world, as unpredictable actions can seriously harm the involved individuals. In this work, we propose a genetic programmin...Votes: 0GitHub stars: 3
- Using Monte Carlo Tree Search As A Demonstrator Within Asynchronous Deep Rl**arXiv ID:** 1812.00045 **Authors:** Bilal Kartal, Pablo Hernandez-Leal, Matthew E. Taylor **Published:** 2018-11-30T20:37:17Z **Abstract:** Deep reinforcement learning (DRL) has achieved great successes in recent years with the help of novel methods and higher compute power. However, there are still several challenges to be addressed such as convergence to locally optimal policies and long training times. In this paper, firstly, we augment Asynchronous Advantage Actor-Critic (A3C) method wi...Votes: 0GitHub stars: 3
- Value Aware Prediction For Robust Multi Agent CoorDerived from arXiv:2607.17914 - Value-Aware Prediction for Robust Multi-Agent Coordination Under Communication LossVotes: 0GitHub stars: 3
- Variance Reduction For Policy Gradient With ActionSkill for AI agent capabilitiesVotes: 0GitHub stars: 3