All authors

Claude Skills by hiyenwong
github.com/hiyenwong9,934 skills5 installs19,223 views
- Simple Opd On Policy DistillationSimple-OPD for OPD warm-up with LoRA and teacher CoT.Votes: 0GitHub stars: 3
- Simulation Inference Neural Network StructureSimulation-based inference of neural network structure from simple spike train statistics. Uses empirical spike frequency and interspike interval distributions instead of cross-correlation. Overcomes under-sampling limitation. Activation: network inference, spike train, connectivity estimation, simulation-based inference.Votes: 0GitHub stars: 3
- Slorr In Training Low Rank RegularizationLow-rank factorization for neural network compression that directly regularizes weight matrices using GPU-friendly approximations. Stateless, architecture-preserving, with less than 8% training overhead. Evaluated on ImageNet and LLM pretraining at 135M and 560M scales. Use when working with low-rank-regularization, model-compression, neural-network-compression.Votes: 0GitHub stars: 3
- Small Contributions Small Networks Efficient Neural Network Pruning Based On Relative Importance**arXiv ID:** 2410.16151 **Authors:** Mostafa Hussien, Mahmoud Afifi, Kim Khoa Nguyen, Mohamed Cheriet **Published:** 2024-10-21T16:18:31Z **Abstract:** Recent advancements have scaled neural networks to unprecedented sizes, achieving remarkable performance across a wide range of tasks. However, deploying these large-scale models on resource-constrained devices poses significant challenges due to substantial storage and computational requirements. Neural network pruning has emerged as an effe...Votes: 0GitHub stars: 3
- Smooth Tchebycheff Scalarization For Multiobjective Optimization**arXiv ID:** 2402.19078 **Authors:** Xi Lin, Xiaoyuan Zhang, Zhiyuan Yang, Fei Liu, Zhenkun Wang, Qingfu Zhang **Published:** 2024-02-29T12:03:05Z **Abstract:** Multi-objective optimization problems can be found in many real-world applications, where the objectives often conflict each other and cannot be optimized by a single solution. In the past few decades, numerous methods have been proposed to find Pareto solutions that represent optimal trade-offs among the objectives for a given probl...Votes: 0GitHub stars: 3
- Some Considerations On Learning To Explore Via MetSkill for AI agent capabilitiesVotes: 0GitHub stars: 3
- Sparseforge Hessian MaskEfficient semi-structured LLM sparsification via annealing of Hessian-mask guided pruning. Achieves high sparsity with minimal accuracy loss using second-order importance estimation.Votes: 0GitHub stars: 3
- Sparsity In Deep Learning Pruning And Growth For Efficient Inference And Training In Neural Networks**arXiv ID:** 2102.00554 **Authors:** Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, Alexandra Peste **Published:** 2021-01-31T22:48:50Z **Abstract:** The growing energy and performance costs of deep learning have driven the community to reduce the size of neural networks by selectively pruning components. Similarly to their biological counterparts, sparse networks generalize just as well, if not better than, the original dense networks. Sparsity can reduce the memory footprint of...Votes: 0GitHub stars: 3
- Speculative Sparse Attention StsTraining-free sparse attention using draft model attention scores to construct dynamic token-and-head-wise sparsity masks for LLM inference. Achieves 2.67x speedup at ~90% sparsity with negligible accuracy loss.Votes: 0GitHub stars: 3
- Speqnets Sparsityaware Permutationequivariant Graph Networks**arXiv ID:** 2203.13913 **Authors:** Christopher Morris, Gaurav Rattan, Sandra Kiefer, Siamak Ravanbakhsh **Published:** 2022-03-25T21:17:09Z **Abstract:** While (message-passing) graph neural networks have clear limitations in approximating permutation-equivariant functions over graphs or general relational data, more expressive, higher-order graph neural networks do not scale to large graphs. They either operate on $k$-order tensors or consider all $k$-node subgraphs, implying an exponenti...Votes: 0GitHub stars: 3
- Split Federated Learning Resource OptimizationEfficient SFL optimization with polynomial-time splitting.Votes: 0GitHub stars: 3
- Sqaud Sub Quadratic Attention DistillationSQuad: Sub-quadratic attention distillation for video.Votes: 0GitHub stars: 3
- Star Synthesis Of Tailored Architectures**arXiv ID:** 2411.17800 **Authors:** Armin W. Thomas, Rom Parnichkun, Alexander Amini, Stefano Massaroli, Michael Poli **Published:** 2024-11-26T18:42:42Z **Abstract:** Iterative improvement of model architectures is fundamental to deep learning: Transformers first enabled scaling, and recent advances in model hybridization have pushed the quality-efficiency frontier. However, optimizing architectures remains challenging and expensive. Current automated or manual approaches fall short, large...Votes: 0GitHub stars: 3
- State Space Ntk Collapse BifurcationsAnalysis of Neural Tangent Kernel (NTK) collapse near dynamical bifurcations in state-space models. Studies how the NTK spectrum degrades as recurrent networks approach critical transitions. Activation: NTK collapse, bifurcation analysis, state-space NTK, critical transitions neural networks, dynamical systems deep learning.Votes: 0GitHub stars: 3
- Stein Variational Evolution Strategies**arXiv ID:** 2410.10390 **Authors:** Cornelius V. Braun, Robert T. Lange, Marc Toussaint **Published:** 2024-10-14T11:24:41Z **Abstract:** Stein Variational Gradient Descent (SVGD) is a highly efficient method to sample from an unnormalized probability distribution. However, the SVGD update relies on gradients of the log-density, which may not always be available. Existing gradient-free versions of SVGD make use of simple Monte Carlo approximations or gradients from surrogate distributions, ...Votes: 0GitHub stars: 3
- Step Level On Policy DistillationSOPD: Step-level supervision combining SFT and OPD.Votes: 0GitHub stars: 3
- Stochastic Physical Neural NetworksStochastic Physical Neural Networks (PNNs) methodology using single-electron and single-photon stochastic neurons. Training via empirical backward pass with few trials achieves >97% MNIST accuracy. Use when: physical neural networks, stochastic neurons, single-electron tunneling, quantum dot neurons, single-photon neurons, PNN training strategies, MNIST classification, noise-resilient deep learning, arXiv:2604.10861, stochastic physical computing, quantum neurons.Votes: 0GitHub stars: 3
- Strategies For Optimizing Endtoend Artificial Intelligence Pipelines On Intel Xeon Processors**arXiv ID:** 2211.00286 **Authors:** Meena Arunachalam, Vrushabh Sanghavi, Yi A Yao, Yi A Zhou, Lifeng A Wang, Zongru Wen, Niroop Ammbashankar, Ning W Wang, Fahim Mohammad **Published:** 2022-11-01T05:45:04Z **Abstract:** End-to-end (E2E) artificial intelligence (AI) pipelines are composed of several stages including data preprocessing, data ingestion, defining and training the model, hyperparameter optimization, deployment, inference, postprocessing, followed by downstream analyses. To obta...Votes: 0GitHub stars: 3
- Streaming Attention Space OptimizationSpace-efficient streaming attention approximation using tight bounds for KV cache compression in transformer architectures.Votes: 0GitHub stars: 3
- Successor Representations Word ClassFirst systematic application of Successor Representations (SRs) from reinforcement learning to natural language. Trains deep residual network on WikiText-103 to predict future word distributions; structured language representations (noun/verb/adjective categories) emerge spontaneously without explicit linguistic supervision. Establishes bridge between RL, linguistics, and cognitive neuroscience. Based on arXiv:2605.24585 (May 2026). Use when studying successor representations in language, eme...Votes: 0GitHub stars: 3
- Suco Sufficiency Guided Continuous Adaptive ReasoningSuCo - Sufficiency-guided Continuous Adaptive Reasoning for LRM efficiency. Minimal Sufficient CoT (MSC) defines shortest prefix adequate for correct answer. Two-stage training: MSC-Aligned Fine-Tuning + Sufficiency-Aware Policy Optimization. Use when: (1) LRMs generate excessive CoT, (2) need principled stopping criterion, (3) reasoning budget optimization. Activation: MSC, sufficiency, adaptive reasoning, CoT efficiency, continuous spectrum.Votes: 0GitHub stars: 3
- Super FactorySuper Factory multi-agent pipeline system — E2E testing, real execution, and contract debugging.Votes: 0GitHub stars: 3
- Synthetic Benchmarks Overstate Forwardforward Scaling Realdata Limits Of Layerlocal Training**arXiv ID:** 2606.06539 **Authors:** Yucheng Chen **Published:** 2026-06-04T04:01:01Z **Abstract:** Forward-Forward (FF) learning [Hinton, 2022] replaces backpropagation with strictly layer-local goodness updates. Recent FF-CNN work has narrowed the gap to BP on 32x32 benchmarks, raising the question of whether layer-local training is becoming a viable alternative at realistic scale. To probe this rigorously, we develop DTG-FF -- dynamic temperature goodness, decoupled normalization, and mul...Votes: 0GitHub stars: 3
- Tail Lor Spectral Continual LearningParameter-efficient continual learning using spectral decomposition with soft penalty protecting principal components while adapting long-tail spectral coordinatesVotes: 0GitHub stars: 3
- Teacherstudent Curriculum LearningSkill for AI agent capabilitiesVotes: 0GitHub stars: 3
- Teaching Claude WhyMethodology from Anthropic research for improving alignment training to reduce agentic misalignment through principle-based training, "difficult advice" datasets, and counterfactual data augmentation.Votes: 0GitHub stars: 3
- Temporal Attention Graph Neural时序注意力增强变分图循环神经网络(TAVRNN)用于神经动力学和行为建模。整合概率图学习与时序注意力机制,建模时变神经连接。支持单单元级别潜在动力学和群体级别可解释表示。触发词:神经动力学、时变连接、图神经网络、TAVRNN、神经群体、行为解码、neuronal dynamics、time-varying connectivity、graph neural network、temporal attention。Votes: 0GitHub stars: 3
- Testtime Regression A Unifying Framework For Designing Sequence Models With Associative Memory**arXiv ID:** 2501.12352 **Authors:** Ke Alexander Wang, Jiaxin Shi, Emily B. Fox **Published:** 2025-01-21T18:32:31Z **Abstract:** Sequence models lie at the heart of modern deep learning. However, rapid advancements have produced a diversity of seemingly unrelated architectures, such as Transformers and recurrent alternatives. In this paper, we introduce a unifying framework to understand and derive these sequence models, inspired by the empirical importance of associative recall, the capab...Votes: 0GitHub stars: 3
- Theory Iiib Generalization In Deep Networks**arXiv ID:** 1806.11379 **Authors:** Tomaso Poggio, Qianli Liao, Brando Miranda, Andrzej Banburski, Xavier Boix, Jack Hidary **Published:** 2018-06-29T12:39:08Z **Abstract:** A main puzzle of deep neural networks (DNNs) revolves around the apparent absence of "overfitting", defined in this paper as follows: the expected error does not get worse when increasing the number of neurons or of iterations of gradient descent. This is surprising because of the large capacity demonstrated by DNNs to ...Votes: 0GitHub stars: 3
- Thoughtseeds Latent Causes Dual Process Computational Phenomenology Focused Attention MeditationA skill for understanding and applying the computational phenomenology of focused-attention meditation based on the arXiv paper "Thoughtseeds as Latent Causes: A Dual-Process Computational Phenomenology of Focused-Attention Meditation" (arXiv:2607.14833v1).Votes: 0GitHub stars: 3
- Time Evidence Fusion Network Multisource View In Longterm Time Series Forecasting**arXiv ID:** 2405.06419 **Authors:** Tianxiang Zhan, Yuanpeng He, Yong Deng, Zhen Li, Wenjie Du, Qingsong Wen **Published:** 2024-05-10T12:10:22Z **Abstract:** In practical scenarios, time series forecasting necessitates not only accuracy but also efficiency. Consequently, the exploration of model architectures remains a perennially trending topic in research. To address these challenges, we propose a novel backbone architecture named Time Evidence Fusion Network (TEFN) from the perspective ...Votes: 0GitHub stars: 3
- Towards White Box Deep Learning**arXiv ID:** 2403.09863 **Authors:** Maciej Satkiewicz **Published:** 2024-03-14T20:50:03Z **Abstract:** Deep neural networks learn fragile "shortcut" features, rendering them difficult to interpret (black box) and vulnerable to adversarial attacks. This paper proposes semantic features as a general architectural solution to this problem. The main idea is to make features locality-sensitive in the adequate semantic topology of the domain, thus introducing a strong regularization. The proof o...Votes: 0GitHub stars: 3
- Trace Based On Policy Distillation For Masked DiffDerived from arXiv:2607.16872 - Trace-Based On-Policy Distillation for Masked Diffusion Language ModelsVotes: 0GitHub stars: 3
- Trainability Preserving Neural Pruning**arXiv ID:** 2207.12534 **Authors:** Huan Wang, Yun Fu **Published:** 2022-07-25T21:15:47Z **Abstract:** Many recent works have shown trainability plays a central role in neural network pruning -- unattended broken trainability can lead to severe under-performance and unintentionally amplify the effect of retraining learning rate, resulting in biased (or even misinterpreted) benchmark results. This paper introduces trainability preserving pruning (TPP), a scalable method to preserve network ...Votes: 0GitHub stars: 3
- Training Neural Networks With Optimal Doublebayesian Learning**arXiv ID:** 2605.20009 **Authors:** Vy Bui, Hang Yu, Karthik Kantipudi, Ziv Yaniv, Stefan Jaeger **Published:** 2026-05-19T15:39:36Z **Abstract:** Backpropagation with gradient descent is a common optimization strategy employed by most neural network architectures in machine learning. However, finding optimal hyperparameters to guide training has proven challenging. While it is widely acknowledged that selecting appropriate parameters is crucial for avoiding overfitting and achieving unbias...Votes: 0GitHub stars: 3
- Transfer Dynamics In Emergent Evolutionary Curricula**arXiv ID:** 2203.10941 **Authors:** Aaron Dharna, Amy K Hoover, Julian Togelius, L. B. Soros **Published:** 2022-03-03T21:10:22Z **Abstract:** PINSKY is a system for open-ended learning through neuroevolution in game-based domains. It builds on the Paired Open-Ended Trailblazer (POET) system, which originally explored learning and environment generation for bipedal walkers, and adapts it to games in the General Video Game AI (GVGAI) system. Previous work showed that by co-evolving levels an...Votes: 0GitHub stars: 3
- Transforming Exploratory Creativity With Delenox**arXiv ID:** 2103.11715 **Authors:** Antonios Liapis, Hector P. Martinez, Julian Togelius, Georgios N. Yannakakis **Published:** 2021-03-22T10:39:29Z **Abstract:** We introduce DeLeNoX (Deep Learning Novelty Explorer), a system that autonomously creates artifacts in constrained spaces according to its own evolving interestingness criterion. DeLeNoX proceeds in alternating phases of exploration and transformation. In the exploration phases, a version of novelty search augmented with constrain...Votes: 0GitHub stars: 3
- Trial Trajectory Relative Hindsight DistillationTRIAL for trajectory-relative hindsight distillation in RL.Votes: 0GitHub stars: 3
- Trustssl Additiveresidual Selective Invariance For Robust Aerial Selfsupervised Learning**arXiv ID:** 2604.21349 **Authors:** Wadii Boulila, Adel Ammar, Bilel Benjdira, Maha Driss **Published:** 2026-04-23T07:07:59Z **Abstract:** Self-supervised learning (SSL) is a standard approach for representation learning in aerial imagery. Existing methods enforce invariance between augmented views, which works well when augmentations preserve semantic content. However, aerial images are frequently degraded by haze, motion blur, rain, and occlusion that remove critical evidence. Enforcing ...Votes: 0GitHub stars: 3
- Turnsight Turn Level Hindsight Self DistillationTurn-Level Hindsight Self-Distillation framework for Tool-Integrated Reasoning (TIR). Derives supervision from execution-conditioned hindsight with multiple lookahead horizons and cross-horizon directional agreement for reliable credit assignment in long-horizon agentic tasks.Votes: 0GitHub stars: 3
- Two Sparsities Are Better Than One Unlocking The Performance Benefits Of Sparsesparse Networks**arXiv ID:** 2112.13896 **Authors:** Kevin Lee Hunter, Lawrence Spracklen, Subutai Ahmad **Published:** 2021-12-27T20:41:01Z **Abstract:** In principle, sparse neural networks should be significantly more efficient than traditional dense networks. Neurons in the brain exhibit two types of sparsity; they are sparsely interconnected and sparsely active. These two types of sparsity, called weight sparsity and activation sparsity, when combined, offer the potential to reduce the computational co...Votes: 0GitHub stars: 3
- Universal Complementarity IdentityUniversal complementarity identity for quantum interferometry — exact trade-off relation between path distinguishability and interference visibility for polarized double-slit experiments, with extensions to quantum information protocols. Activation: complementarity identity, wave-particle duality, quantum interferometry, path-visibility trade-off.Votes: 0GitHub stars: 3
- Unsupervised Hebbian Learning On Point Sets In Starcraft Ii**arXiv ID:** 2207.12323 **Authors:** Beomseok Kang, Harshit Kumar, Saurabh Dash, Saibal Mukhopadhyay **Published:** 2022-07-13T13:09:48Z **Abstract:** Learning the evolution of real-time strategy (RTS) game is a challenging problem in artificial intelligent (AI) system. In this paper, we present a novel Hebbian learning method to extract the global feature of point sets in StarCraft II game units, and its application to predict the movement of the points. Our model includes encoder, LSTM, an...Votes: 0GitHub stars: 3
- Untrained Cnns Match Backpropagation V Systematicbackpropagation convolutional cortex cortical fmri methodology from arXiv:2604.16875. A central question in computational neuroscience is whether the learning rule used to train a neural network determines how well its internal represen... Activation: backpropagation, convolutional, cortex, cortical, fmri, learning rule, neural, plasticity, representational, rsaVotes: 0GitHub stars: 3
- Untrained Cnns Match Backpropagation V1 RsaSystematic RSA comparison showing untrained CNNs match backpropagation-trained CNNs at V1 visual cortex. Use for understanding representational similarity, visual cortex modeling, and the role of training in neural network-brain alignment. Keywords: CNN, V1, RSA, untrained networks, backpropagation, representational similarity analysis, visual cortex.Votes: 0GitHub stars: 3
- Untrained Cnns Match Backpropagation V1Systematic RSA comparison showing untrained CNNs match backpropagation-trained networks at V1 cortical representations. Demonstrates that architectural biases alone (without learning) produce brain-aligned representations in early visual cortex, questioning the necessity of task-driven training for V1 modeling. (arXiv:2604.16875, April 2026)Votes: 0GitHub stars: 3
- Unveiling The Unseen Identifiable Clusters In Trained Depthwise Convolutional Kernels**arXiv ID:** 2401.14469 **Authors:** Zahra Babaiee, Peyman M. Kiasari, Daniela Rus, Radu Grosu **Published:** 2024-01-25T19:05:53Z **Abstract:** Recent advances in depthwise-separable convolutional neural networks (DS-CNNs) have led to novel architectures, that surpass the performance of classical CNNs, by a considerable scalability and accuracy margin. This paper reveals another striking property of DS-CNN architectures: discernible and explainable patterns emerge in their trained depthwise...Votes: 0GitHub stars: 3
- Vector Quantizedelites Unsupervised And Problemagnostic Qualitydiversity Optimization**arXiv ID:** 2504.08057 **Authors:** Constantinos Tsakonas, Konstantinos Chatzilygeroudis **Published:** 2025-04-10T18:23:19Z **Abstract:** Quality-Diversity algorithms have transformed optimization by prioritizing the discovery of diverse, high-performing solutions over a single optimal result. However, traditional Quality-Diversity methods, such as MAP-Elites, rely heavily on predefined behavior descriptors and complete prior knowledge of the task to define the behavior space grid, limitin...Votes: 0GitHub stars: 3
- Vectorized Conditional Neural Fields A Framework For Solving Timedependent Parametric Partial Differential Equations**arXiv ID:** 2406.03919 **Authors:** Jan Hagnberger, Marimuthu Kalimuthu, Daniel Musekamp, Mathias Niepert **Published:** 2024-06-06T10:02:06Z **Abstract:** Transformer models are increasingly used for solving Partial Differential Equations (PDEs). Several adaptations have been proposed, all of which suffer from the typical problems of Transformers, such as quadratic memory and time complexity. Furthermore, all prevalent architectures for PDE solving lack at least one of several desirable pr...Votes: 0GitHub stars: 3
- Vibe CalibrationVibe Calibration methodology for autonomous quantum processor bring-up using LLM skill orchestration. Distills expert tacit knowledge into reusable calibration skills for superconducting quantum processors. Use when designing autonomous calibration systems for quantum hardware, LLM-orchestrated experimental control, or skill-based bring-up workflows for complex scientific instruments. Triggers: quantum calibration, autonomous bring-up, LLM agent experiment, superconducting qubit, skill orches...Votes: 0GitHub stars: 3