All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,376
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 12,625–12,648 of 21,376 skills
- Vista Path An Interactive Foundation Model For PatImplement techniques from VISTA-PATH: An interactive foundation model for pathology image segmentation and quantitative analysis in computational pathology. Accurate semantic segmentation for histopathology image is crucial for quantitative tissue analysis and downstream clinical modelingVotes: 0GitHub stars: 6
- Visgym Diverse Customizable Scalable EnvironmentsImplement techniques from VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents. The suite spans symbolic puzzles, real-image understanding, navigation, and manipulation, and provides flexible controls over difficulty, input representation, planning horizon, and feedbackVotes: 0GitHub stars: 6
- Video Deep Research AgentImplements video deep research for multi-hop reasoning combining video analysis, web search, and evidence synthesis. Evaluates workflow vs agentic paradigms with 100-sample benchmark across 6 semantic domains, revealing goal drift and long-horizon consistency as core bottlenecks.Votes: 0GitHub stars: 6
- Verse Embedding Visualization DocumentsOptimize vision-language models for document tasks via embedding visualization and clustering-guided data generation. Identify error-prone regions in visual space and synthetically augment training data targeting weak areas.Votes: 0GitHub stars: 6
- Verse Craft Video World ModelsControl video generation via 4D geometric representation combining static background point clouds and per-object 3D Gaussian trajectories. Enable category-agnostic control over camera and multi-object motion in realistic video synthesis.Votes: 0GitHub stars: 6
- Uniworld Semantic VisionCombine semantic encoders from multimodal LLMs with contrastive learning to create unified high-resolution encoders for both visual understanding and generation tasks without relying on VAE compression.Votes: 0GitHub stars: 6
- Typhoon Ocr Open Vision Language Model For ThaiDocument extraction is a core component of digital workflows, yet existing vision-language models (VLMs) predominantly favor high-resource languages. Thai presents additional challenges due to script complexity from non-latin letters, the absence of explicit word boundaries, and the prevalence of highly unstructured real-world documents, limiting the effectiveness of current open-source models. This paper presents Typhoon OCR, an open VLM for document extraction tailored for Thai and English....Votes: 0GitHub stars: 6
- Twinbrainvla Unleashing The Potential Of GeneralisImplement techniques from TwinBrainVLA: Unleashing the Potential of Generalist VLMs for Embodied Tasks via Asymmetric Mixture-of-Transformers. The fundamental premise of Vision-Language-Action (VLA) models is to harness the extensive general capabilities of pre-trained Vision-Language Models (VLMs) for generalized embodied intelligenceVotes: 0GitHub stars: 6
- Towards Pixel Level Vlm Perception Via Simple PoinImplement techniques from Towards Pixel-Level VLM Perception via Simple Points Prediction. We present SimpleSeg, a strikingly simple yet highly effective approach to endow Multimodal Large Language Models (MLLMs) with native pixel-level perceptionVotes: 0GitHub stars: 6
- Towards Efficient And Robust Linguistic EmotionLinguistic expressions of emotions such as depression, anxiety, and trauma-related states are pervasive in clinical notes, counseling dialogues, and online mental health communities, and accurate recognition of these emotions is essential for clinical triage, risk assessment, and timely intervention. Although large language models (LLMs) have demonstrated strong generalization ability in emotion analysis tasks, their diagnostic reliability in high-stakes, context-intensive medical settings re...Votes: 0GitHub stars: 6
- Toward Efficient Agents Memory Tool Learning AndRecent years have witnessed increasing interest in extending large language models into agentic systems. While the effectiveness of agents has continued to improve, efficiency, which is crucial for real-world deployment, has often been overlooked. This paper therefore investigates efficiency from three core components of agents: memory, tool learning, and planning, considering costs such as latency, tokens, steps, etc. Aimed at conducting comprehensive research addressing the efficiency of th...Votes: 0GitHub stars: 6
- Toolprmbench Evaluating And Advancing ProcessReward-guided search methods have demonstrated strong potential in enhancing tool-using agents by effectively guiding sampling and exploration over complex action spaces. As a core design, those search methods utilize process reward models (PRMs) to provide step-level rewards, enabling more fine-grained monitoring. However, there is a lack of systematic and reliable evaluation benchmarks for PRMs in tool-using settings. In this paper, we introduce ToolPRMBench, a large-scale benchmark specifi...Votes: 0GitHub stars: 6
- Tool Verification ReasoningTool Verification stabilizes self-improving reasoning models by using external tool execution as ground-truth evidence to prevent spurious consensus from becoming reinforced training signals.Votes: 0GitHub stars: 6
- The Responsibility Vacuum Organizational FailureModern CI/CD pipelines integrating agent-generated code exhibit a structural failure in responsibility attribution. Decisions are executed through formally correct approval processes, yet no entity possesses both the authority to approve those decisions and the epistemic capacity to meaningfully understand their basis. We define this condition as responsibility vacuum: a state in which decisions occur, but responsibility cannot be attributed because authority and verification capacity do not ...Votes: 0GitHub stars: 6
- Standing Committee MoeReveal that Mixture-of-Experts models harbor a 'Standing Committee' of consistent expert coalitions handling majority computational load across domains. Challenges specialization assumptions and suggests training approaches like load-balancing losses may work against natural optimization.Votes: 0GitHub stars: 6
- Sparsemm Visual AttentionDiscovers that <5% of attention heads process visual information in MLLMs; introduces SparseMM for asymmetric KV-cache allocation achieving 1.38x acceleration and 52% memory reduction.Votes: 0GitHub stars: 6
- Sliding Window Attention AdaptationAdapt full-attention language models to sliding window attention without expensive retraining. Combine five synergistic strategies (full decode, sink tokens, interleaved layers, chain-of-thought, fine-tuning) achieving 30-100% speedups while maintaining 90-100% accuracy.Votes: 0GitHub stars: 6
- Sin Bench Tracing Native Evidence Chains In LongEvaluating whether multimodal large language models truly understand long-form scientific papers remains challenging: answer-only metrics and synthetic 'Needle-In-A-Haystack' tests often reward answer matching without requiring a causal, evidence-linked reasoning trace in the document. We propose the 'Fish-in-the-Ocean' (FITO) paradigm, which requires models to construct explicit cross-modal evidence chains within native scientific documents. To operationalize FITO, we build SIN-Data, a scien...Votes: 0GitHub stars: 6
- Simplevla Rl Scaling Vla TrainingApply reinforcement learning to Vision-Language-Action models for robotic control, achieving 99% LIBERO task success and discovering novel manipulation strategies (pushcut) without task-specific reward engineering. Scales efficiently via parallelized trajectory sampling and outcome-based rewards.Votes: 0GitHub stars: 6
- Scientific Image Synthesis Benchmarking MethodologImplement techniques from Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility. While synthetic data has proven effective for improving scientific reasoning in the text domain, multimodal reasoning remains constrained by the difficulty of synthesizing scientifically rigorous imagesVotes: 0GitHub stars: 6
- Sciarena Evaluation PlatformBuild community-driven evaluation platforms for scientific tasks using pairwise model comparisons and human voting. Assess foundation models on literature-grounded reasoning without automated metrics.Votes: 0GitHub stars: 6
- Scaling Behavior Cloning GamesTrain video game-playing foundation models discovering that increasing training data and network depth enables learning more causal policies. Release 8300+ hours of gameplay data and open-source models for real-time consumer GPU inference.Votes: 0GitHub stars: 6
- Salad Achieve High Sparsity Attention Via EfficienImplement techniques from SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer. Diffusion Transformers have recently demonstrated remarkable performance in video generationVotes: 0GitHub stars: 6
- Robovip Robot Video SynthesisGenerate synthetic robot manipulation data via diffusion models using visual identity prompting from exemplar images. Improve multi-view temporal coherence and scalability for robot policy training without extensive real-world data collection.Votes: 0GitHub stars: 6