All authors

Claude Skills by hiyenwong
github.com/hiyenwong9,934 skills5 installs19,223 views
- Multi Robot Rigidity ControlAngle-based localization and rigidity maintenance control for multi-robot networks under sensing constraints. Establishes equivalence between angle rigidity and bearing rigidity with directed sensing graphs and body-frame bearing measurements. Use for: multi-robot formation control, angle-based localization, rigidity maintenance, bearing rigidity analysis, decentralized robot control. Activation: multi-robot rigidity, angle-based localization, bearing rigidity, formation control, rigidity mai...Votes: 0GitHub stars: 3
- Multiobjective Modelbased Policy Search For Dataefficient Learning With Sparse Rewards**arXiv ID:** 1806.09351 **Authors:** Rituraj Kaushik, Konstantinos Chatzilygeroudis, Jean-Baptiste Mouret **Published:** 2018-06-25T09:46:47Z **Abstract:** The most data-efficient algorithms for reinforcement learning in robotics are model-based policy search algorithms, which alternate between learning a dynamical model of the robot and optimizing a policy to maximize the expected return given the model and its uncertainties. However, the current algorithms lack an effective exploration str...Votes: 0GitHub stars: 3
- Mute Communication Unlearning Multi AgentReturn-preserving communication unlearning for efficient multi-agent coordination under bandwidth constraints. Enables MARL agents to selectively forget inter-agent communications while preserving coordination returns, addressing the trade-off between communication sparsity and cooperative performance. Activation: MUTE, communication unlearning, MARL, bandwidth constraints, cooperative games, inter-agent communication, partial observability.Votes: 0GitHub stars: 3
- Neo Agentic Program AnalysisNeo agentic program analysis framework for detecting privilege escalation in polyglot microservices. Combines LLM-based agents with classic program analysis for cross-service vulnerability detection across multiple languages. Uses dynamic analysis planning, adaptive code search primitives, and semantic validation. Use when: analyzing microservice security, detecting privilege escalation, performing agentic code analysis across polyglot codebases, building LLM-assisted program analysis tools, ...Votes: 0GitHub stars: 3
- Neural Architecture Search With Reinforcement Learning**arXiv ID:** 1611.01578 **Authors:** Barret Zoph, Quoc V. Le **Published:** 2016-11-05T00:41:37Z **Abstract:** Neural networks are powerful and flexible models that work well for many difficult learning tasks in image, speech and natural language understanding. Despite their success, neural networks are still hard to design. In this paper, we use a recurrent network to generate the model descriptions of neural networks and train this RNN with reinforcement learning to maximize the expected a...Votes: 0GitHub stars: 3
- Neurorefiner Multi Agent Neuron SegmentationNeuroRefiner multi-agent 3D neuron segmentation framework.Votes: 0GitHub stars: 3
- Nl Cps Reinforcement Learning Based Kubernetes ConThe placement of Kubernetes control-plane nodes is critical to ensuring cluster reliability, scalability, and performance, and therefore represents a significant deployment challen... Activation: reinforcement, learning, based, kubernetes, controlVotes: 0GitHub stars: 3
- Nl Cps Reinforcement Learning Based Kubernetes ControlThe placement of Kubernetes control-plane nodes is critical to ensuring cluster reliability, scalability, and performance, and therefore represents a significant deployment challenge in heterogeneous,... Activation: kubernetes control plane, multi-region cluster, RL-based orchestration, reconfigurable intelligent surfaces, RIS.Votes: 0GitHub stars: 3
- Nl Cps Reinforcement Learning Based KubernetesThe placement of Kubernetes control-plane nodes is critical to ensuring cluster reliability, scalability, and performance, and therefore represents a ... Activation: control, optimal control, MPC, distributed systems, multi-agentVotes: 0GitHub stars: 3
- Nonmarkovian Control With Gated Endtoend Memory Policy Networks**arXiv ID:** 1705.10993 **Authors:** Julien Perez, Tomi Silander **Published:** 2017-05-31T09:00:44Z **Abstract:** Partially observable environments present an important open challenge in the domain of sequential control learning with delayed rewards. Despite numerous attempts during the two last decades, the majority of reinforcement learning algorithms and associated approximate models, applied to this context, still assume Markovian state transitions. In this paper, we explore the use of ...Votes: 0GitHub stars: 3
- Npas A Compileraware Framework Of Unified Network Pruning And Architecture Search For Beyond Realtime Mobile Acceleration**arXiv ID:** 2012.00596 **Authors:** Zhengang Li, Geng Yuan, Wei Niu, Pu Zhao, Yanyu Li, Yuxuan Cai, Xuan Shen, Zheng Zhan, Zhenglun Kong, Qing Jin, Zhiyu Chen, Sijia Liu, Kaiyuan Yang, Bin Ren, Yanzhi Wang, Xue Lin **Published:** 2020-12-01T16:03:40Z **Abstract:** With the increasing demand to efficiently deploy DNNs on mobile edge devices, it becomes much more important to reduce unnecessary computation and increase the execution speed. Prior methods towards this goal, including model comp...Votes: 0GitHub stars: 3
- On The Convergence And Stability Of Upsidedown Reinforcement Learning Goalconditioned Supervised Learning And Online Decision Transformers**arXiv ID:** 2502.05672 **Authors:** Miroslav Štrupl, Oleg Szehr, Francesco Faccio, Dylan R. Ashley, Rupesh Kumar Srivastava, Jürgen Schmidhuber **Published:** 2025-02-08T19:26:22Z **Abstract:** This article provides a rigorous analysis of convergence and stability of Episodic Upside-Down Reinforcement Learning, Goal-Conditioned Supervised Learning and Online Decision Transformers. These algorithms performed competitively across various benchmarks, from games to robotic tasks, but their theo...Votes: 0GitHub stars: 3
- On The Importance Of Critical Period In Multistage Reinforcement Learning**arXiv ID:** 2208.04832 **Authors:** Junseok Park, Inwoo Hwang, Min Whoo Lee, Hyunseok Oh, Minsu Lee, Youngki Lee, Byoung-Tak Zhang **Published:** 2022-08-09T15:17:22Z **Abstract:** The initial years of an infant's life are known as the critical period, during which the overall development of learning performance is significantly impacted due to neural plasticity. In recent studies, an AI agent, with a deep neural network mimicking mechanisms of actual neurons, exhibited a learning period si...Votes: 0GitHub stars: 3
- On The Power Of Gradual Network Alignment Using Dualperception Similarities**arXiv ID:** 2201.10945 **Authors:** Jin-Duk Park, Cong Tran, Won-Yong Shin, Xin Cao **Published:** 2022-01-26T14:01:32Z **Abstract:** Network alignment (NA) is the task of finding the correspondence of nodes between two networks based on the network structure and node attributes. Our study is motivated by the fact that, since most of existing NA methods have attempted to discover all node pairs at once, they do not harness information enriched through interim discovery of node correspondenc...Votes: 0GitHub stars: 3
- One Shot Imitation LearningSkill for AI agent capabilitiesVotes: 0GitHub stars: 3
- Openforgerl Train Harness Native Agents In Any EnvironmentOpenForgeRL: Train Harness-native Agents in Any EnvironmentVotes: 0GitHub stars: 3
- Optimization Trilemma Efficiency Comfort Fairness Decentralized Multi Agent CoordinatioSkill derived from arXiv:2607.17311 - The Optimization Trilemma: Efficiency, Comfort and Fairness in Decentralized Multi-agent CoordinatioVotes: 0GitHub stars: 3
- Oracle Multi Objective Rl Circuit DesignMulti-objective reinforcement learning framework for analog circuit design optimization using LLM-guided exploration and preference-aware conditioning.Votes: 0GitHub stars: 3
- Organizational Memory Agentic Business ProcessLLM-based agents for automating business process execution using organizational memory. General-purpose LLMs lack organization-specific knowledge; this work extracts and structures organizational knowledge from human-oriented artifacts to enable reliable agentic process execution. Activation: organizational memory, business process automation, agentic bpm, knowledge extraction, process execution, LLM agents.Votes: 0GitHub stars: 3
- Otap Structure Aware Optimal Transport Evaluating Planning Execution Agent TrajectoriesSkill derived from arXiv:2607.17082 - Otap:Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent TrajectoriesVotes: 0GitHub stars: 3
- Pac Bench Privacy Multi AgentResearch paper: PAC-BENCH - Evaluating Multi-Agent Collaboration under Privacy Constraints. First benchmark for differential privacy in multi-agent systems.Votes: 0GitHub stars: 3
- Parallel Exploration Via Negatively Correlated Search**arXiv ID:** 1910.07151 **Authors:** Peng Yang, Qi Yang, Ke Tang, Xin Yao **Published:** 2019-10-16T03:30:29Z **Abstract:** Effective exploration is a key to successful search. The recently proposed Negatively Correlated Search (NCS) tries to achieve this by parallel exploration, where a set of search processes are driven to be negatively correlated so that different promising areas of the search space can be visited simultaneously. Various applications have verified the advantages of such n...Votes: 0GitHub stars: 3
- Particle Swarm Optimization For Generating Interpretable Fuzzy Reinforcement Learning Policies**arXiv ID:** 1610.05984 **Authors:** Daniel Hein, Alexander Hentschel, Thomas Runkler, Steffen Udluft **Published:** 2016-10-19T12:41:52Z **Abstract:** Fuzzy controllers are efficient and interpretable system controllers for continuous state and action spaces. To date, such controllers have been constructed manually or trained automatically either using expert-generated problem-specific cost functions or incorporating detailed knowledge about the optimal control strategy. Both requirements f...Votes: 0GitHub stars: 3
- Partition Logic Social ComplementarityPartition-logic framework for modeling complementarity in social measurement. Applies non-Boolean event structures (partition logics) to social-science settings where mutually incompatible observation modes reveal different aspects of a definite latent state. Use when: social complementarity, measurement incompatibility in social science, partition logic applications, quantum-inspired social measurement, personnel assessment modeling, survey design, organizational auditing.Votes: 0GitHub stars: 3
- Pats Policy Aware Training Scaffolding For Agentic ReinforcementPATS: Policy-Aware Training Scaffolding for Agentic Reinforcement LearningVotes: 0GitHub stars: 3
- Paving Way Agents BiologyMethodology from Anthropic research (Jun 2026) on making biological data infrastructure agent-friendly. Case study shows that adding deterministic retrieval layers (like gget virus) to scientific research agents improves accuracy from inconsistent results to nearly 100% for dataset construction tasks.Votes: 0GitHub stars: 3
- Pdeflow Autonomous Agentic Pde PipelinesPDEFlow: an autonomous agentic framework that turns user-level ODE/PDE descriptions into solver-backed neural-operator pipelines. Links problem specification, data generation, operator training, and checkpoint-based inference via a stateful input graph and registry-based interface. Instantiated with multi-branch Bayesian DeepONet. Activation: PDEFlow, autonomous PDE solver, neural operator, agentic pipeline, DeepONet, FEniCSx, ODE PDE automation, scientific workflow, Bayesian DeepONet, operat...Votes: 0GitHub stars: 3
- Pearl Auditable Repair For Scientific Reasoning GrDerived from arXiv:2607.17917 - PEARL: Auditable Repair for Scientific Reasoning Graph ExtractionVotes: 0GitHub stars: 3
- Pearl Auditable Repair Scientific Reasoning Graph ExtractionSkill derived from arXiv:2607.17917 - PEARL: Auditable Repair for Scientific Reasoning Graph ExtractionVotes: 0GitHub stars: 3
- Perfagent Profiler Guided Iterative Refinement ForSkill generated from arXiv paper 2607.19653: PerfAgent: Profiler-Guided Iterative Refinement for Repository-Level Code OptimizationVotes: 0GitHub stars: 3
- Performing Deep Recurrent Double Qlearning For Atari Games**arXiv ID:** 1908.06040 **Authors:** Felipe Moreno-Vera **Published:** 2019-08-16T15:56:16Z **Abstract:** Currently, many applications in Machine Learning are based on define new models to extract more information about data, In this case Deep Reinforcement Learning with the most common application in video games like Atari, Mario, and others causes an impact in how to computers can learning by himself with only information called rewards obtained from any action. There is a lot of algorithm...Votes: 0GitHub stars: 3
- Physics Audited Agentic Discovery In Scientific Machine LearningIn agentic scientific machine learning (SciML), large language model (LLM) agents can discover surrogate models and select one by an automated score, typically an error metric. A low error, however, d. Based on arXiv:2607.07379.Votes: 0GitHub stars: 3
- Physics Aware Quadcopter Drl ControlPhysics-aware end-to-end deep reinforcement learning methodology for quadcopter control with actuator dynamics modeling.Votes: 0GitHub stars: 3
- Pisas Contextual Integrity Multi User AgenticBenchmarking contextual integrity in multi-user agentic systems. As LLM agents evolve into shared organizational infrastructure, new privacy risks emerge from inter-agent messages, shared memory, and cross-user information exposure. Activation: contextual integrity, multi-user agents, privacy benchmark, agentic privacy, inter-agent communication, shared memory privacy.Votes: 0GitHub stars: 3
- Playing The Lottery With Rewards And Multiple Languages Lottery Tickets In Rl And Nlp**arXiv ID:** 1906.02768 **Authors:** Haonan Yu, Sergey Edunov, Yuandong Tian, Ari S. Morcos **Published:** 2019-06-06T18:38:38Z **Abstract:** The lottery ticket hypothesis proposes that over-parameterization of deep neural networks (DNNs) aids training by increasing the probability of a "lucky" sub-network initialization being present rather than by helping the optimization process (Frankle & Carbin, 2019). Intriguingly, this phenomenon suggests that initialization strategies for DNNs can be...Votes: 0GitHub stars: 3
- Position Leverage Foundational Models For Blackbox Optimization**arXiv ID:** 2405.03547 **Authors:** Xingyou Song, Yingtao Tian, Robert Tjarko Lange, Chansoo Lee, Yujin Tang, Yutian Chen **Published:** 2024-05-06T15:10:46Z **Abstract:** Undeniably, Large Language Models (LLMs) have stirred an extraordinary wave of innovation in the machine learning research domain, resulting in substantial impact across diverse fields such as reinforcement learning, robotics, and computer vision. Their incorporation has been rapid and transformative, marking a significan...Votes: 0GitHub stars: 3
- Possibility Before Utility Learning And Using Hierarchical Affordances**arXiv ID:** 2203.12686 **Authors:** Robby Costales, Shariq Iqbal, Fei Sha **Published:** 2022-03-23T19:17:22Z **Abstract:** Reinforcement learning algorithms struggle on tasks with complex hierarchical dependency structures. Humans and other intelligent agents do not waste time assessing the utility of every high-level action in existence, but instead only consider ones they deem possible in the first place. By focusing only on what is feasible, or "afforded", at the present moment, an agen...Votes: 0GitHub stars: 3
- Predictive Coding Graphs Are A Superset Of Feedforward Neural Networks**arXiv ID:** 2603.06142 **Authors:** Björn van Zwol **Published:** 2026-03-06T10:50:41Z **Abstract:** Predictive coding graphs (PCGs) are a recently introduced generalization to predictive coding networks, a neuroscience-inspired probabilistic latent variable model. Here, we prove how PCGs define a mathematical superset of feedforward artificial neural networks (multilayer perceptrons). This positions PCNs more strongly within contemporary machine learning (ML), and reinforces earlier propos...Votes: 0GitHub stars: 3
- Prime Plasticity Recovery In Multi Agent EnvironmeDerived from arXiv:2607.17922 - PRIME: Plasticity Recovery in Multi-Agent Environments for UAV-Assisted Emergency Communication NetworksVotes: 0GitHub stars: 3
- Proactive Memory Agent Long Horizon TasksProactive memory agent that runs alongside an unmodified action agent to prevent behavioral state decay in long-horizon tasks. Updates structured memory bank from trajectory and selectively injects reminders. Plug-and-play with frontier agents. +8.3pp on Terminal-Bench 2.0, +6.8pp on τ²-Bench. Activation: proactive memory, long-horizon agents, trajectory management, context surfacing, memory retrieval, behavioral state decay.Votes: 0GitHub stars: 3
- Program Synthesis Through Reinforcement Learning Guided Tree Search**arXiv ID:** 1806.02932 **Authors:** Riley Simmons-Edler, Anders Miltner, Sebastian Seung **Published:** 2018-06-08T00:53:43Z **Abstract:** Program Synthesis is the task of generating a program from a provided specification. Traditionally, this has been treated as a search problem by the programming languages (PL) community and more recently as a supervised learning problem by the machine learning community. Here, we propose a third approach, representing the task of synthesizing a given pro...Votes: 0GitHub stars: 3
- Proof Carrying Agent ActionsRuntime governance framework for heterogeneous agent systems using action certificates. Model-agnostic governance centered on action proofs rather than vendor-native session records. From Anthropic research (arXiv:2606.04104).Votes: 0GitHub stars: 3
- Proximal Policy Optimization With Evolutionary Mutations**arXiv ID:** 2601.14705 **Authors:** Casimir Czworkowski, Stephen Hornish, Alhassan S. Yasin **Published:** 2026-01-21T06:34:53Z **Abstract:** Proximal Policy Optimization (PPO) is a widely used reinforcement learning algorithm known for its stability and sample efficiency, but it often suffers from premature convergence due to limited exploration. In this paper, we propose POEM (Proximal Policy Optimization with Evolutionary Mutations), a novel modification to PPO that introduces an adaptiv...Votes: 0GitHub stars: 3
- Qftuner Breaking Tradition In Reinforcement Learning**arXiv ID:** 2402.16562 **Authors:** Mahmood A. Jumaah, Yossra H. Ali, Tarik A. Rashid **Published:** 2024-02-26T13:39:04Z **Abstract:** In reinforcement learning algorithms, the hyperparameters tuning method refers to choosing the optimal parameters that may increase the overall performance. Manual or random hyperparameter tuning methods can lead to different results in the reinforcement learning algorithms. In this paper, we propose a new method called QF-tuner for automatic hyperparameter...Votes: 0GitHub stars: 3
- Qmarl Entanglement CoordinationQuantum Multi-Agent Reinforcement Learning (QMARL) with entanglement-based coordination. Demonstrates provable quantum advantage via CHSH game (Tsirelson limit 0.854 vs classical ceiling 0.75). Hybrid quantum actor + classical critic outperforms both fully classical and fully quantum.Votes: 0GitHub stars: 3
- Quantifying Generalization In Reinforcement LearniSkill for AI agent capabilitiesVotes: 0GitHub stars: 3
- Quantum Deep Q Learning With Distributed Prioritized Experience Replay**arXiv ID:** 2304.09648 **Authors:** Samuel Yen-Chi Chen **Published:** 2023-04-19T13:40:44Z **Abstract:** This paper introduces the QDQN-DPER framework to enhance the efficiency of quantum reinforcement learning (QRL) in solving sequential decision tasks. The framework incorporates prioritized experience replay and asynchronous training into the training algorithm to reduce the high sampling complexities. Numerical simulations demonstrate that QDQN-DPER outperforms the baseline distributed ...Votes: 0GitHub stars: 3
- Quctrl Bell CompilerCompiler-driven sub-microsecond feedback control stack for trapped-ion quantum experiments. Use when designing quantum control software stacks, compiler pipelines for hardware control, deterministic low-latency feedback systems, DSL transpilation, or real-time quantum hardware synchronization.Votes: 0GitHub stars: 3
- Racing Control Variable Genetic Programming For Symbolic Regression**arXiv ID:** 2309.07934 **Authors:** Nan Jiang, Yexiang Xue **Published:** 2023-09-13T21:38:06Z **Abstract:** Symbolic regression, as one of the most crucial tasks in AI for science, discovers governing equations from experimental data. Popular approaches based on genetic programming, Monte Carlo tree search, or deep reinforcement learning learn symbolic regression from a fixed dataset. They require massive datasets and long training time especially when learning complex equations involving ...Votes: 0GitHub stars: 3
- Rapid Tasksolving In Novel Environments**arXiv ID:** 2006.03662 **Authors:** Sam Ritter, Ryan Faulkner, Laurent Sartran, Adam Santoro, Matt Botvinick, David Raposo **Published:** 2020-06-05T20:09:20Z **Abstract:** We propose the challenge of rapid task-solving in novel environments (RTS), wherein an agent must solve a series of tasks as rapidly as possible in an unfamiliar environment. An effective RTS agent must balance between exploring the unfamiliar environment and solving its current task, all while building a model of the ne...Votes: 0GitHub stars: 3