All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,402
- 892
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 10,393–10,416 of 21,402 skills
- Recurrent Neural Networks With External Memory For Language Understanding**arXiv ID:** 1506.00195 **Authors:** Baolin Peng, Kaisheng Yao **Published:** 2015-05-31T05:10:03Z **Abstract:** Recurrent Neural Networks (RNNs) have become increasingly popular for the task of language understanding. In this task, a semantic tagger is deployed to associate a semantic label to each word in an input sequence. The success of RNN may be attributed to its ability to memorize long-term dependence that relates the current-time semantic label prediction to the observations many ti...Votes: 0GitHub stars: 3
- Multidimensional Cv Qkd ReconciliationMultidimensional reconciliation methodology for continuous-variable QKD with HDirac open-source simulation frameworkVotes: 0GitHub stars: 3
- Filtering Variational Objectives**arXiv ID:** 1705.09279 **Authors:** Chris J. Maddison, Dieterich Lawson, George Tucker, Nicolas Heess, Mohammad Norouzi, Andriy Mnih, Arnaud Doucet, Yee Whye Teh **Published:** 2017-05-25T17:52:41Z **Abstract:** When used as a surrogate objective for maximum likelihood estimation in latent variable models, the evidence lower bound (ELBO) produces state-of-the-art results. Inspired by this, we consider the extension of the ELBO to a family of lower bounds defined by a particle filter's estim...Votes: 0GitHub stars: 3
- Certify Ed Multi Layer Verification Framework ExactExact diagonalization (ED) is a workhorse technique in computational quantum many-body physics, but published ED results are rarely accompanied by machine-checkable evidence of their numerical correctVotes: 0GitHub stars: 3
- Arxiv 2608 18988v1 Deepweaver Bridging The Evidence Synthesis Gap In**arXiv ID:** 2608.18988v1 **Authors:** Xujia Wang, Yizhe Zhang, Bin Xu, Lei Hou, Juanzi Li **URL:** http://arxiv.org/abs/2608.18988v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- A Twostage Approach To Devicerobust Acoustic Scene Classification**arXiv ID:** 2011.01447 **Authors:** Hu Hu, Chao-Han Huck Yang, Xianjun Xia, Xue Bai, Xin Tang, Yajian Wang, Shutong Niu, Li Chai, Juanjuan Li, Hongning Zhu, Feng Bao, Yuanjun Zhao, Sabato Marco Siniscalchi, Yannan Wang, Jun Du, Chin-Hui Lee **Published:** 2020-11-03T03:27:18Z **Abstract:** To improve device robustness, a highly desirable key feature of a competitive data-driven acoustic scene classification (ASC) system, a novel two-stage system based on fully convolutional neural networks ...Votes: 0GitHub stars: 3
- A Deep Neural Network Surrogate Modeling Benchmark For Temperature Field Prediction Of Heat Source Layout**arXiv ID:** 2103.11177 **Authors:** Xianqi Chen, Xiaoyu Zhao, Zhiqiang Gong, Jun Zhang, Weien Zhou, Xiaoqian Chen, Wen Yao **Published:** 2021-03-20T13:26:21Z **Abstract:** Thermal issue is of great importance during layout design of heat source components in systems engineering, especially for high functional-density products. Thermal analysis generally needs complex simulation, which leads to an unaffordable computational burden to layout optimization as it iteratively evaluates different...Votes: 0GitHub stars: 3
- V2a Cross Domain Offline RlV2A methodology — unifying Value Alignment, Assignment, and dynamics alignment for cross-domain offline RL with heterogeneous datasets from multiple source domains collected by diverse behavior policies.Votes: 0GitHub stars: 3
- Ttrl Cocov Test Time Rl ConfidenceTest-Time Reinforcement Learning with Confidence-Conditioned Verification (TTRL-CoCoV) methodology for optimizing Pass@k coverage and Pass@1 performance in label-free settings.Votes: 0GitHub stars: 3
- Som Score Based Meanflow Policy OptimizationSOM (Score-Based One-step MeanFlow Policy Optimization) — actor-critic algorithm combining MeanFlow with online RL using score estimation and probability flow ODE.Votes: 0GitHub stars: 3
- Sbsrl Sampling Based Safe RlSBSRL — Sampling-based safe RL with joint constraint enforcement across dynamics samples and epistemic uncertainty exploration constraints.Votes: 0GitHub stars: 3
- Reward Uncertainty Diverse BehaviourReformulate RL objective using reward function distribution instead of scalar reward. Apply non-linear objective over action sets to induce calibrated behavioural diversity without sacrificing expected reward.Votes: 0GitHub stars: 3
- Rat Randomized Advantage TransformationRandomized Advantage Transformation (RAT) methodology for computing Tikhonov-regularized natural policy gradients via direct backpropagation. Uses Woodbury formula and randomized block Kaczmarz iterations to avoid explicit Fisher matrix construction, CG solvers, or architecture-specific approximations. ICML 2026 accepted. Matches or exceeds established natural-gradient methods across continuous and visual control benchmarks. Use when: scalable natural policy gradients, Fisher-free natural gra...Votes: 0GitHub stars: 3
- Precise Sde Consistent Rl Flow MatchingPrecise — SDE-consistent stochastic sampling for RL post-training of flow-matching models with clean-latent posterior mean freezing.Votes: 0GitHub stars: 3
- Lilac Safe Continual RlLILAC+ — Safe continual RL under nonstationarity with adaptive safety constraints (context-based, adaptation-speed, budget-to-state).Votes: 0GitHub stars: 3
- Kl Trajectory Decoupling Llm DistillationKL-Trajectory Decoupling methodology — unified theoretical framework decomposing LLM distillation into two orthogonal choices: prefix distribution (what to condition on) and trajectory distribution (how to generate responses). Reveals that SFT, DAgger, Offline RL, and On-Policy Distillation (OPD) differ along these two axes. Use when: analyzing distillation methods, choosing between SFT/Dagger/OPD/Offline-RL, designing new distillation algorithms, understanding KL divergence in LLM fine-tunin...Votes: 0GitHub stars: 3
- Efficient TdmpcEfficientTDMPC improves model-based RL for continuous control with ensemble dynamics, uncertainty-penalized planning, and data freshness optimizations. Achieves SOTA sample efficiency on HumanoidBench-Hard and DMC hard, with benefits from higher update-to-data ratios.Votes: 0GitHub stars: 3
- Delta Discriminative Token Credit AssignmentDelTA (Discriminative Token Credit Assignment) methodology for Reinforcement Learning from Verifiable Rewards (RLVR). Introduces a discriminator view of RLVR updates showing policy-gradient implicitly acts as a linear discriminator over token-gradient vectors. Proposes token-level coefficient estimation to amplify discriminative directions and downweight shared patterns (e.g. formatting tokens). Outperforms baselines by 3.26 pts on Qwen3-8B and 2.62 pts on Qwen3-14B across math benchmarks. Us...Votes: 0GitHub stars: 3
- Clipping Bottleneck NsrNear-boundary Stochastic Rescue (NSR) for stabilizing RLVR/GRPO training via stochastic recovery of clipped signalsVotes: 0GitHub stars: 3
- Vla Probabilistic Chunk MaskingDrop-in GRPO modification that allocates gradient computation to a small, probabilistically selected subset of trajectory chunks using success-failure action variance. Achieves 2.38x wall-clock speedup while matching final performance.Votes: 0GitHub stars: 3
- Stepwise Reasoning SubgraphStepwise reasoning framework that builds query-specific subgraphs from external knowledge bases to ground intermediate reasoning steps, improving LLM reasoning accuracy and factual reliability.Votes: 0GitHub stars: 3
- Reasoning Driven RetrievalRetrieval as iterative reasoning methodology. Treat retrieval as explicit hypothesis-driven search with evidence evaluation and self-improving refinement. Use when building RAG systems, information retrieval agents, search optimization, or any system that needs to go beyond black-box retrieval to find latent-pattern documents.Votes: 0GitHub stars: 3
- Nonlinear Cross Entropy BenchmarkingSample-efficient quantum advantage benchmarking using nonlinear cross-entropy and heavy output generation classifiers. Use when: (1) benchmarking NISQ quantum circuits, (2) distinguishing quantum computers from classical spoofers, (3) designing quantum advantage experiments, (4) analyzing random circuit sampling results, (5) evaluating shallow-depth quantum circuits.Votes: 0GitHub stars: 3
- Modular State Space Model 260714078Model human perception, cognition, and decision dynamics as a modular perception-cognition-decision pipeline state-space model. Provides mathematical formulation, stability conditions, and application to rehabilitation control. Use when you need interpretable dynamical models linking neural mechanisms to behavior.Votes: 0GitHub stars: 3