All authors

Claude Skills by hiyenwong
github.com/hiyenwong9,934 skills5 installs19,223 views
- Gcpo Cooperative Policy OptimizationGroup Cooperative Policy Optimization (GCPO) replaces GRPO's winner-takes-all competition with team cooperation. Rollouts are rewarded by contribution to team's valid solution coverage, measured as determinant volume over reward-weighted semantic embeddings. Solves exploration collapse in RLVR for LLM reasoning.Votes: 0GitHub stars: 3
- Kl Trajectory Decoupling Llm DistillationKL-Trajectory Decoupling methodology — unified theoretical framework decomposing LLM distillation into two orthogonal choices: prefix distribution (what to condition on) and trajectory distribution (how to generate responses). Reveals that SFT, DAgger, Offline RL, and On-Policy Distillation (OPD) differ along these two axes. Use when: analyzing distillation methods, choosing between SFT/Dagger/OPD/Offline-RL, designing new distillation algorithms, understanding KL divergence in LLM fine-tunin...Votes: 0GitHub stars: 3
- Learning Zone Energy Data SelectionLearning-Zone Energy methodology — online data selection for efficient RL post-training of LLMs. Identifies the 'learning zone' where samples have optimal difficulty for gradient signal, replacing uniform rollout/gradient budgets in GRPO/DAPO. Use when: optimizing RL post-training compute, data selection for LLM RL, GRPO efficiency, DAPO optimization, RL training data prioritization, reasoning model post-training. Activation: learning zone energy, online data selection RL, GRPO data efficienc...Votes: 0GitHub stars: 3
- Lilac Safe Continual RlLILAC+ — Safe continual RL under nonstationarity with adaptive safety constraints (context-based, adaptation-speed, budget-to-state).Votes: 0GitHub stars: 3
- Linear Recurrent Memory Pomdp RlTheoretical justification for why linear recurrent neural networks work as memory units in partially observable RL. Constructs linear filters that reproduce HMM belief logits and achieve vanishing state-decoding error under near-deterministic transitions.Votes: 0GitHub stars: 3
- Llm Sleep Paradigm Self Modify ConsolidateSleep-Dreaming paradigm for LLM continual learning via Knowledge Seeding (on-policy distillation + RL imitation learning) and RL-generated synthetic curriculum for self-improvement.Votes: 0GitHub stars: 3
- Magic Multi Step MarlMAGIC (Multi-step Advantage-Gated Interventional Causal MARL) — counterfactual action interventions with advantage-gated intrinsic rewards for multi-agent coordination.Votes: 0GitHub stars: 3
- Medgym Continuous Time Medical RlMedGym - Unified continuous-time benchmark for dynamic medical treatment reinforcement learning using Physics-Informed Neural NetworksVotes: 0GitHub stars: 3
- Model Based Diffusion Policy OptimizationModel-Based Diffusion Policy Optimization (MBDPO) methodology for scaling world-model reinforcement learning. Unifies search and policy optimization through diffusion policy representations addressing structural misalignment. Use for world-model RL, diffusion-based policy learning, offline pretraining, model-based RL scaling.Votes: 0GitHub stars: 3
- Oppo Token Credit AssignmentOracle-Prompted Policy Optimization (OPPO) for token-level credit assignment in LLM reasoning via Bayesian value recursionVotes: 0GitHub stars: 3
- Pg Dpo Non Exponential DiscountingPontryagin-Guided Direct Policy Optimization (PG-DPO) — a variational RL framework that replaces Bellman recursions with Pontryagin Maximum Principle for non-exponential discounting (hyperbolic, survival-discount). Handles settings where standard value/actor-critic methods fail. Use when: RL with non-exponential discounting, hyperbolic discounting RL, human-like time preferences, survival processes, Pontryagin-based RL. Activation: PG-DPO, non-exponential discounting, Pontryagin RL, hyperboli...Votes: 0GitHub stars: 3
- Precise Sde Consistent Rl Flow MatchingPrecise — SDE-consistent stochastic sampling for RL post-training of flow-matching models with clean-latent posterior mean freezing.Votes: 0GitHub stars: 3
- Qldpc Full Extractor ConstructionFull extractor construction for logical processing in Hypergraph Product (HGP) codes — surgery systems for measuring arbitrary logical Pauli operators on QLDPC code blocks. Enables Pauli-based computation without compilation overhead. Use when: QLDPC code processing, logical operator measurement, hypergraph product codes, fault-tolerant quantum memory, quantum error correction, Pauli-based computation.Votes: 0GitHub stars: 3
- Rat Randomized Advantage TransformationRandomized Advantage Transformation (RAT) methodology for computing Tikhonov-regularized natural policy gradients via direct backpropagation. Uses Woodbury formula and randomized block Kaczmarz iterations to avoid explicit Fisher matrix construction, CG solvers, or architecture-specific approximations. ICML 2026 accepted. Matches or exceeds established natural-gradient methods across continuous and visual control benchmarks. Use when: scalable natural policy gradients, Fisher-free natural gra...Votes: 0GitHub stars: 3
- Research Paper Pattern ExtractorExtract reusable research skill patterns from knowledge graph paper analysis. Uses PageRank, Louvain, and vector search to identify important papers and research clusters, then distills patterns into new skills. Activation: extract research pattern, research skill extractor, 研究模式提炼, paper pattern analysis.Votes: 0GitHub stars: 3
- Reward Uncertainty Diverse BehaviourReformulate RL objective using reward function distribution instead of scalar reward. Apply non-linear objective over action sets to induce calibrated behavioural diversity without sacrificing expected reward.Votes: 0GitHub stars: 3
- Sbsrl Sampling Based Safe RlSBSRL — Sampling-based safe RL with joint constraint enforcement across dynamics samples and epistemic uncertainty exploration constraints.Votes: 0GitHub stars: 3
- Selfplay Data Gating CollapseAnalysis of data gating vs reward grounding in self-play RL for LLMs, revealing the Grounded Proposer Paradox and two-stage phase transitionsVotes: 0GitHub stars: 3
- Skill Rm Unifying Reward ModelingSkill Reward Model (Skill-RM) framework for unified reward modeling in RL pipelines. Treats reward computation as structured agentic task orchestrating heterogeneous evaluation criteria.Votes: 0GitHub stars: 3
- Som Score Based Meanflow Policy OptimizationSOM (Score-Based One-step MeanFlow Policy Optimization) — actor-critic algorithm combining MeanFlow with online RL using score estimation and probability flow ODE.Votes: 0GitHub stars: 3
- Survival Reinforcement LearningSurvival Reinforcement Learning (SRL) - online classification-based self-supervised RL that maximizes agent dwell time at target goals, extending survival value learning framework. Bypasses contrastive RL constraints and mitigates "bang-bang" control issues.Votes: 0GitHub stars: 3
- Ttrl Cocov Test Time Rl ConfidenceTest-Time Reinforcement Learning with Confidence-Conditioned Verification (TTRL-CoCoV) methodology for optimizing Pass@k coverage and Pass@1 performance in label-free settings.Votes: 0GitHub stars: 3
- Two Timescale Markovian Sa ConvergenceConvergence theory for two-timescale stochastic approximations under Markovian noise, applicable to TDC and actor-critic methods. First almost sure convergence proof for TDC with eligibility traces under off-policy learning with linear function approximation. ICML 2026.Votes: 0GitHub stars: 3
- V2a Cross Domain Offline RlV2A methodology — unifying Value Alignment, Assignment, and dynamics alignment for cross-domain offline RL with heterogeneous datasets from multiple source domains collected by diverse behavior policies.Votes: 0GitHub stars: 3
- Runtime Monitoring Distributed Cps No Global ClockMonitor distributed CPS without global clock using STL.Votes: 0GitHub stars: 3
- Scalable Training Continuous Time Snn DstdScalable Training of Continuous-Time Spiking Neural Networks with Differentiable Spike-Time Discretization (DSTD) — reduces memory and training time for deep SNNs via fixed-time discretization and synfire-chain-inspired regularizationVotes: 0GitHub stars: 3
- Scalable Training Continuous Time Spiking Neural Networks DstdThis methodology addresses the severe memory constraints in training deep continuous-time Spiking Neural Networks (SNNs). Traditional exact spike-time computation requires evaluating and retaining candidate firing times over intervals determined by presynaptic spike ordering, leading to memory complexity of O(N_out * N_in) which becomes prohibitive for deep networks.Votes: 0GitHub stars: 3
- A Bandit Approach With Evolutionary Operators For Model Selection**arXiv ID:** 2402.05144 **Authors:** Margaux Brégère, Julie Keisler **Published:** 2024-02-07T08:01:45Z **Abstract:** This work formulates model selection as an infinite-armed bandit problem, namely, a problem in which a decision maker iteratively selects one of an infinite number of fixed choices (i.e., arms) when the properties of each choice are only partially known at the time of allocation and may become better understood over time, via the attainment of rewards.Here, the arms are machi...Votes: 0GitHub stars: 3
- A Deep Neural Network Surrogate Modeling Benchmark For Temperature Field Prediction Of Heat Source Layout**arXiv ID:** 2103.11177 **Authors:** Xianqi Chen, Xiaoyu Zhao, Zhiqiang Gong, Jun Zhang, Weien Zhou, Xiaoqian Chen, Wen Yao **Published:** 2021-03-20T13:26:21Z **Abstract:** Thermal issue is of great importance during layout design of heat source components in systems engineering, especially for high functional-density products. Thermal analysis generally needs complex simulation, which leads to an unaffordable computational burden to layout optimization as it iteratively evaluates different...Votes: 0GitHub stars: 3
- A Firefly Algorithm For Mixedvariable Optimization Based On Hybrid Distance Modeling**arXiv ID:** 2603.26792 **Authors:** Ousmane Tom Bechir, Adán José-García, Zaineb Chelly Garcia, Vincent Sobanski, Clarisse Dhaenens **Published:** 2026-03-25T16:52:54Z **Abstract:** Several real-world optimization problems involve mixed-variable search spaces, where continuous, ordinal, and categorical decision variables coexist. However, most population-based metaheuristic algorithms are designed for either continuous or discrete optimization problems and do not naturally handle heterogene...Votes: 0GitHub stars: 3
- A Generative Neural Annealer For Blackbox Combinatorial Optimization**arXiv ID:** 2505.09742 **Authors:** Yuan-Hang Zhang, Massimiliano Di Ventra **Published:** 2025-05-14T19:05:19Z **Abstract:** We propose a generative, end-to-end solver for black-box combinatorial optimization that emphasizes both sample efficiency and solution quality on NP problems. Drawing inspiration from annealing-based algorithms, we treat the black-box objective as an energy function and train a neural network to model the associated Boltzmann distribution. By conditioning on tempera...Votes: 0GitHub stars: 3
- A High Speed Multilabel Classifier Based On Extreme Learning Machines**arXiv ID:** 1608.08898 **Authors:** Meng Joo Er, Rajasekar Venkatesan, Ning Wang **Published:** 2016-08-31T14:56:12Z **Abstract:** In this paper a high speed neural network classifier based on extreme learning machines for multi-label classification problem is proposed and dis-cussed. Multi-label classification is a superset of traditional binary and multi-class classification problems. The proposed work extends the extreme learning machine technique to adapt to the multi-label problems. As...Votes: 0GitHub stars: 3
- A Learningbased Cooperative Coevolution Framework For Heterogeneous Largescale Global Optimization**arXiv ID:** 2604.01241 **Authors:** Wenjie Qiu, Zixin Wang, Hongyu Fang, Zeyuan Ma, Yue-Jiao Gong **Published:** 2026-03-30T00:18:05Z **Abstract:** Cooperative Coevolution (CC) effectively addresses Large-Scale Global Optimization (LSGO) via decomposition but struggles with the emerging class of Heterogeneous LSGO (H-LSGO) problems arising from real-world applications, where subproblems exhibit diverse dimensions and distinct landscapes. The prevailing CC paradigm, relying on a fixed low-di...Votes: 0GitHub stars: 3
- A Novel Metaheuristic Optimization Algorithm Inspired By The Spread Of Viruses**arXiv ID:** 2006.06282 **Authors:** Zhixi Li, Vincent Tam **Published:** 2020-06-11T09:35:28Z **Abstract:** According to the no-free-lunch theorem, there is no single meta-heuristic algorithm that can optimally solve all optimization problems. This motivates many researchers to continuously develop new optimization algorithms. In this paper, a novel nature-inspired meta-heuristic optimization algorithm called virus spread optimization (VSO) is proposed. VSO loosely mimics the spread of viru...Votes: 0GitHub stars: 3
- A Novel Online Realtime Classifier For Multilabel Data Streams**arXiv ID:** 1608.08905 **Authors:** Rajasekar Venkatesan, Meng Joo Er, Shiqian Wu, Mahardhika Pratama **Published:** 2016-08-31T15:14:06Z **Abstract:** In this paper, a novel extreme learning machine based online multi-label classifier for real-time data streams is proposed. Multi-label classification is one of the actively researched machine learning paradigm that has gained much attention in the recent years due to its rapidly increasing real world applications. In contrast to traditional...Votes: 0GitHub stars: 3
- A Survey On Reservoir Computing And Its Interdisciplinary Applications Beyond Traditional Machine Learning**arXiv ID:** 2307.15092 **Authors:** Heng Zhang, Danilo Vasconcellos Vargas **Published:** 2023-07-27T05:20:20Z **Abstract:** Reservoir computing (RC), first applied to temporal signal processing, is a recurrent neural network in which neurons are randomly connected. Once initialized, the connection strengths remain unchanged. Such a simple structure turns RC into a non-linear dynamical system that maps low-dimensional inputs into a high-dimensional space. The model's rich dynamics, linear s...Votes: 0GitHub stars: 3
- A Twostage Approach To Devicerobust Acoustic Scene Classification**arXiv ID:** 2011.01447 **Authors:** Hu Hu, Chao-Han Huck Yang, Xianjun Xia, Xue Bai, Xin Tang, Yajian Wang, Shutong Niu, Li Chai, Juanjuan Li, Hongning Zhu, Feng Bao, Yuanjun Zhao, Sabato Marco Siniscalchi, Yannan Wang, Jun Du, Chin-Hui Lee **Published:** 2020-11-03T03:27:18Z **Abstract:** To improve device robustness, a highly desirable key feature of a competitive data-driven acoustic scene classification (ASC) system, a novel two-stage system based on fully convolutional neural networks ...Votes: 0GitHub stars: 3
- A Unified Detection Framework For Ai Related Content And ArtifactsArtificial intelligence (AI) is a double-edged sword: while it has achieved remarkable success across a wide range of domains, its deployment also calls for effective oversight and regulation, for whi. Based on arXiv:2607.07527.Votes: 0GitHub stars: 3
- Active Learning For Neural Pde Solvers**arXiv ID:** 2408.01536 **Authors:** Daniel Musekamp, Marimuthu Kalimuthu, David Holzmüller, Makoto Takamoto, Mathias Niepert **Published:** 2024-08-02T18:48:58Z **Abstract:** Solving partial differential equations (PDEs) is a fundamental problem in science and engineering. While neural PDE solvers can be more efficient than established numerical solvers, they often require large amounts of training data that is costly to obtain. Active learning (AL) could help surrogate models reach the sam...Votes: 0GitHub stars: 3
- Active2 Learning Actively Reducing Redundancies In Active Learning Methods For Sequence Tagging And Machine Translation**arXiv ID:** 2103.06490 **Authors:** Rishi Hazra, Parag Dutta, Shubham Gupta, Mohammed Abdul Qaathir, Ambedkar Dukkipati **Published:** 2021-03-11T06:27:31Z **Abstract:** While deep learning is a powerful tool for natural language processing (NLP) problems, successful solutions to these problems rely heavily on large amounts of annotated samples. However, manually annotating data is expensive and time-consuming. Active Learning (AL) strategies reduce the need for huge volumes of labeled data...Votes: 0GitHub stars: 3
- Adnev The Multiarchitecture Neuroevolutionbased Multivariate Anomaly Detection Framework**arXiv ID:** 2404.07968 **Authors:** Marcin Pietroń, Dominik Żurek, Kamil Faber, Roberto Corizzo **Published:** 2024-03-25T08:40:58Z **Abstract:** Anomaly detection tools and methods enable key analytical capabilities in modern cyberphysical and sensor-based systems. Despite the fast-paced development in deep learning architectures for anomaly detection, model optimization for a given dataset is a cumbersome and time-consuming process. Neuroevolution could be an effective and efficient solut...Votes: 0GitHub stars: 3
- Agentic Sabre Ransomware DetectionUncertainty-aware neuro-symbolic multi-agent framework for adaptive ransomware detection. Fuses semantic representation evidence with behavioural forensic telemetry, using Monte Carlo Dropout for epistemic uncertainty quantification. A risk-uncertainty orchestrator triages cases: auto-contain high-confidence threats, escalate uncertain cases to humans. Includes post-hoc explainability (gradient saliency, permutation importance, counterfactual analysis). Activation: agentic SABRE, ransomware d...Votes: 0GitHub stars: 3
- Aggregate In The Advantage Not The Ratio A CanonicDerived from arXiv:2607.17924 - Aggregate in the Advantage, Not the Ratio: A Canonical-Form Analysis of Cooperative Multi-Agent Policy OptimizationVotes: 0GitHub stars: 3
- Aigasdevl An Adaptive Incremental Neural Gas Model For Drifting Data Streams Under Extreme Verification Latency**arXiv ID:** 2407.05379 **Authors:** Maria Arostegi, Miren Nekane Bilbao, Jesus L. Lobo, Javier Del Ser **Published:** 2024-07-07T14:04:57Z **Abstract:** The ever-growing speed at which data are generated nowadays, together with the substantial cost of labeling processes cause Machine Learning models to face scenarios in which data are partially labeled. The extreme case where such a supervision is indefinitely unavailable is referred to as extreme verification latency. On the other hand, in...Votes: 0GitHub stars: 3
- Alphaevolve A Coding Agent For Scientific And Algorithmic Discovery**arXiv ID:** 2506.13131 **Authors:** Alexander Novikov, Ngân Vũ, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco J. R. Ruiz, Abbas Mehrabian, M. Pawan Kumar, Abigail See, Swarat Chaudhuri, George Holland, Alex Davies, Sebastian Nowozin, Pushmeet Kohli, Matej Balog **Published:** 2025-06-16T06:37:18Z **Abstract:** In this white paper, we present AlphaEvolve, an evolutionary coding agent that substantially enhances capabili...Votes: 0GitHub stars: 3
- Alphaschema Exploring The Space Of Trading SemantiAlphaSchema: Exploring the Space of Trading Semantics for LLM-Based Alpha MiningVotes: 0GitHub stars: 3
- An Efficient Nonconvex Reformulation Of Stagewise Convex Optimization Problems**arXiv ID:** 2010.14322 **Authors:** Rudy Bunel, Oliver Hinder, Srinadh Bhojanapalli, Krishnamurthy, Dvijotham **Published:** 2020-10-27T14:30:32Z **Abstract:** Convex optimization problems with staged structure appear in several contexts, including optimal control, verification of deep neural networks, and isotonic regression. Off-the-shelf solvers can solve these problems but may scale poorly. We develop a nonconvex reformulation designed to exploit this staged structure. Our reformulati...Votes: 0GitHub stars: 3
- An Online Universal Classifier For Binary Multiclass And Multilabel Classification**arXiv ID:** 1609.00843 **Authors:** Meng Joo Er, Rajasekar Venkatesan, Ning Wang **Published:** 2016-09-03T17:03:14Z **Abstract:** Classification involves the learning of the mapping function that associates input samples to corresponding target label. There are two major categories of classification problems: Single-label classification and Multi-label classification. Traditional binary and multi-class classifications are sub-categories of single-label classification. Several classifiers a...Votes: 0GitHub stars: 3
- Architect Regularize And Replay Arr A Flexible Hybrid Approach For Continual Learning**arXiv ID:** 2301.02464 **Authors:** Vincenzo Lomonaco, Lorenzo Pellegrini, Gabriele Graffieti, Davide Maltoni **Published:** 2023-01-06T11:22:59Z **Abstract:** In recent years we have witnessed a renewed interest in machine learning methodologies, especially for deep representation learning, that could overcome basic i.i.d. assumptions and tackle non-stationary environments subject to various distributional shifts or sample selection biases. Within this context, several computational approa...Votes: 0GitHub stars: 3
- Arxiv 2608 05436 The Ethics Of Artificial Intelligence In The LifeThe ethics of artificial intelligence in the life sciences: Universality, cultural diversity and an architecture of care (arXiv: 2608.05436)Votes: 0GitHub stars: 3