All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,376
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 12,649–12,672 of 21,376 skills
- Rehy At Video Diffusion AttentionMerge softmax and linear attention for video diffusion models using chunk-wise recurrent reformulation with constant memory usage. Enable efficient distillation from existing softmax models, reducing training cost two orders of magnitude to ~160 GPU hours.Votes: 0GitHub stars: 6
- Rebuttalagent Strategic Persuasion In Academic RebImplement techniques from RebuttalAgent: Strategic Persuasion in Academic Rebuttal via Theory of Mind. Although artificial intelligence (AI) has become deeply integrated into various stages of the research workflow and achieved remarkable advancements, academic rebuttal remains a significant and underexplored challengeVotes: 0GitHub stars: 6
- Re Align Reasoning Image GenerationBridge image understanding-generation gap via In-Context Chain-of-Thought reasoning and RL training with surrogate rewards. Improve faithful execution of mixed image-text prompts in generation and editing tasks.Votes: 0GitHub stars: 6
- Pyramidal Wan Video EfficiencyConvert pretrained video diffusion models into pyramidal architectures via low-cost finetuning while preserving output quality. Explore step distillation for enhanced efficiency, enabling deployment of efficient inference without training from scratch.Votes: 0GitHub stars: 6
- Prorl Reasoning ExpansionTrain LLMs to discover novel reasoning strategies beyond base model capabilities using prolonged RL with KL control, reference policy resets, and diverse task suites.Votes: 0GitHub stars: 6
- Profuse 3d Semantic UnderstandingApply semantic understanding to 3D Gaussian Splatting scenes through dense correspondence-guided pre-registration without render-supervised fine-tuning. Achieve semantic understanding in ~5 minutes using cross-view clustering and direct language feature fusion.Votes: 0GitHub stars: 6
- Privileged Information Object DetectionLeverage training-time privileged information (depth, saliency maps) to improve student detector performance without inference overhead. Model-agnostic methodology applicable across detection architectures with no increase in inference complexity.Votes: 0GitHub stars: 6
- Prism Benchmarking Phone Realization In SpeechPhone recognition (PR) serves as the atomic interface for language-agnostic modeling for cross-lingual speech processing and phonetic analysis. Despite prolonged efforts in developing PR systems, current evaluations only measure surface-level transcription accuracy. We introduce PRiSM, the first open-source benchmark designed to expose blind spots in phonetic perception through intrinsic and extrinsic evaluation of PR systems. PRiSM standardizes transcription-based evaluation and assesses dow...Votes: 0GitHub stars: 6
- Pretraining Midtraining Rl InterplayUnderstand when RL genuinely expands reasoning beyond pre-training through controlled experiments on synthetic tasks. Discover that RL works best at the edge of competence and process rewards reduce hacking—critical for designing effective reasoning model training.Votes: 0GitHub stars: 6
- Plenoptic Video GenerationGenerate spatially and temporally coherent multi-view video through autoregressive conditioning with camera-guided retrieval and progressive context scaling. Enable long-video generation maintaining spatio-temporal memory across viewpoint changes.Votes: 0GitHub stars: 6
- Physrvg Physics Aware Unified ReinforcementPhysical principles are fundamental to realistic visual simulation, but remain a significant oversight in transformer-based video generation. This gap highlights a critical limitation in rendering rigid body motion, a core tenet of classical mechanics. While computer graphics and physics-based simulators can easily model such collisions using Newton formulas, modern pretrain-finetune paradigms discard the concept of object rigidity during pixel-level global denoising. Even perfectly correct m...Votes: 0GitHub stars: 6
- Opennovelty Scholarly AssessmentBuild agentic systems for transparent, evidence-based novelty analysis of research submissions through four-phase pipelines: contribution extraction, prior work retrieval, hierarchical comparison, and structured reporting with explicit citations—enabling fair peer review at scale.Votes: 0GitHub stars: 6
- One Sample Polymath LearningDemonstrate that a single strategically engineered training sample can improve reasoning across multiple domains. Polymath learning shows sample quality and multidisciplinary design matter more than quantity, enabling extreme data efficiency in RL training.Votes: 0GitHub stars: 6
- Multi Scale Speculative DecodingAccelerate autoregressive image generation via multi-resolution drafting with spatially-informed verification. Local rejection and resampling enable efficient error correction focusing on spatial neighborhoods, achieving 1.7× speedup over baselines.Votes: 0GitHub stars: 6
- Molecular Thought ReasoningImprove agent reasoning by designing thought structures that balance deep analysis, self-reflection, and exploratory thinking. Framework discovers that effective long-form reasoning exhibits molecular-like interaction patterns—specific bonds between reasoning components that enable fast entropy convergence. Method synthesizes improved reasoning trajectories using distribution-transfer, improving both model performance and RL training stability.Votes: 0GitHub stars: 6
- Memoryrewardbench Benchmarking Reward Models ForExisting works increasingly adopt memory-centric mechanisms to process long contexts in a segment manner, and effective memory management is one of the key capabilities that enables large language models to effectively propagate information across the entire sequence. Therefore, leveraging reward models (RMs) to automatically and reliably evaluate memory quality is critical. In this work, we introduce MemoryRewardBench, the first benchmark to systematically study the ability of RMs to evaluat...Votes: 0GitHub stars: 6
- Memorization 3d Shape GenerationEvaluate memorization in 3D generative models through controlled experiments discovering factors like dataset diversity and guidance scale. Provide simple yet effective strategies like rotation augmentation to reduce memorization without degrading generation quality.Votes: 0GitHub stars: 6
- Meeplelm A Virtual Playtester Simulating Diverse SImplement techniques from MeepleLM: A Virtual Playtester Simulating Diverse Subjective Experiences. Recent advancements have expanded the role of Large Language Models in board games from playing agents to creative co-designersVotes: 0GitHub stars: 6
- Medical Video GenerationGenerate accurate, high-quality medical videos for clinical education and documentation by leveraging large-scale annotated medical datasets with domain-specific fine-tuning on video diffusion models.Votes: 0GitHub stars: 6
- Medical Sam3 A Foundation Model For UniversalPromptable segmentation foundation models such as SAM3 have demonstrated strong generalization capabilities through interactive and concept-based prompting. However, their direct applicability to medical image segmentation remains limited by severe domain shifts, the absence of privileged spatial prompts, and the need to reason over complex anatomical and volumetric structures. Here we present Medical SAM3, a foundation model for universal prompt-driven medical image segmentation, obtained by...Votes: 0GitHub stars: 6
- Mecellem Models Turkish Models Trained From ScratcImplement techniques from Mecellem Models: Turkish Models Trained from Scratch and Continually Pre-trained for the Legal Domain. This paper presents Mecellem models, a framework for developing specialized language models for the Turkish legal domain through domain adaptation strategiesVotes: 0GitHub stars: 6
- Longcat Flash Thinking 2601 Technical ReportImplement techniques from LongCat-Flash-Thinking-2601 Technical Report. LongCat-Flash-Thinking-2601 achieves state-of-the-art performance among open-source models on a wide range of agentic benchmarks, including agentic search, agentic tool use, and tool-integrated reasoningVotes: 0GitHub stars: 6
- Locate Steer And Improve A Practical Survey OfMechanistic Interpretability (MI) has emerged as a vital approach to demystify the opaque decision-making of Large Language Models (LLMs). However, existing reviews primarily treat MI as an observational science, summarizing analytical insights while lacking a systematic framework for actionable intervention. To bridge this gap, we present a practical survey structured around the pipeline: 'Locate, Steer, and Improve.' We formally categorize Localizing (diagnosis) and Steering (intervention) ...Votes: 0GitHub stars: 6
- Liberty A Causal Framework For BenchmarkingConcept-based explanations quantify how high-level concepts (e.g., gender or experience) influence model behavior, which is crucial for decision-makers in high-stakes domains. Recent work evaluates the faithfulness of such explanations by comparing them to reference causal effects estimated from counterfactuals. In practice, existing benchmarks rely on costly human-written counterfactuals that serve as an imperfect proxy. To address this, we introduce a framework for constructing datasets con...Votes: 0GitHub stars: 6