Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,839
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,081–7,104 of 20,839 skills

Habitat Predictor EvalA

Probes the ability of a runtime-based predictor to accurately estimate GPU training iteration execution times and cost-normalized throughput across different DNN architectures and GPU generations without requiring full training runs. Use when the user wants to benchmark on ImageNet, WMT'16, LSUN, or asks about evaluating this task. Reports average prediction error.

researchpythonperformance
0
3
Habitat Gs EvalA

Evaluates embodied agents' navigation capabilities in photorealistic 3D Gaussian Splatting environments versus traditional mesh-based simulators. It probes cross-domain generalization, visual robustness, and human-aware collision avoidance in dynamic scenes. Use when the user wants to benchmark on InteriorGS + Real-world GS, Habitat-Matterport 3D (HM3D), AnimatableGaussians, or asks about evaluating this task. Reports SR, SPL.

researchpythongo
0
3
Habitat Benchmark EvalA

Evaluates a robot's ability to perform long-horizon mobile manipulation and object rearrangement tasks in simulated environments. It probes hierarchical planning, whole-body continuous control, and robust recovery from failures across multi-step subtask sequences. Use when the user wants to benchmark on Habitat Benchmark, or asks about evaluating this task. Reports completion rate.

researchpythongo
0
3
Habibi Tts EvalA

Evaluates zero-shot and unified-dialectal text-to-speech synthesis across multiple Arabic dialects. It measures transcription accuracy, speaker similarity, and audio naturalness to assess how well a model preserves dialectal features and voice identity without dialect-specific fine-tuning. Use when the user wants to benchmark on Habibi Benchmark, or asks about evaluating this task. Reports WER-O.

researchpython
0
3
H3wb EvalA

Evaluates 3D whole-body human pose estimation and lifting capabilities. It probes a model's ability to reconstruct 133-keypoint 3D skeletons from complete 2D poses, occluded/incomplete 2D poses, or monocular RGB images, with specific focus on body, face, and hand regions. Use when the user wants to benchmark on H3WB, or asks about evaluating this task. Reports MPJPE.

researchpythongit
0
3
H2vu Benchmark EvalA

Evaluates multimodal large language models on hierarchical and holistic video understanding, specifically probing temporal reasoning, countercommonsense comprehension, trajectory state tracking, and first-person streaming video analysis. Use when the user wants to benchmark on H²VU, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
H2seqrec EvalA

Evaluates sequential recommendation models on predicting the next item a user will interact with, capturing temporal dynamics and handling sparse user-item interactions. Use when the user wants to benchmark on AMT, Goodreads, or asks about evaluating this task. Reports HR@K, NDCG@K.

researchpythongo
0
3
Gym V EvalA

Evaluates agentic vision models on zero-shot generalization across 179 procedurally generated environments spanning 10 domains. It measures task completion via answer correctness for single-turn interactions and cumulative performance via normalized episodic return for multi-turn interactions. Use when the user wants to benchmark on Gym-V, or asks about evaluating this task. Reports normalized episodic return.

researchpythongo
0
3
Gwtc3 Spin Population EvalA

Evaluates the ability of population inference models to explain the observed spin distributions of binary black holes in the GWTC-3 catalog. It probes how well different astrophysical formation scenarios (e.g., field vs. dynamical assembly, zero-spin subpopulations) fit the gravitational-wave data. Use when the user wants to benchmark on GWTC-3, or asks about evaluating this task. Reports Bayes factor ($\mathcal{B}$).

researchpythongit
0
3
Gwtc2 Bbh Mass Distribution EvalA

Evaluates the ability of semi-parametric and parametric models to recover the astrophysical primary mass distribution of binary black holes from gravitational wave observations, specifically testing for features like the ~35 M⊙ peak and low-mass structure. Use when the user wants to benchmark on GWTC-2 catalog, or asks about evaluating this task. Reports marginal likelihood.

researchpythongo
0
3
Gwrfrf Spatial Spectrum EvalA

This benchmark evaluates a model's ability to synthesize accurate spatial radio-frequency spectra at target transmitter locations using neighboring spectra and scene geometry. It probes both single-scene prediction accuracy and cross-scene generalization capabilities in wireless propagation environments. Use when the user wants to benchmark on RFID Dataset, MATLAB Dataset, or asks about evaluating this task. Reports MSE.

researchpython
0
3
Gwlans EvalA

Predicts target words in computer-aided translation based on source sentences, translation context (prefix, suffix, zero, bidirectional), and human-typed characters. It probes the model's ability to handle discontinuous context and weak positional information in real-world CAT scenarios. Use when the user wants to benchmark on GWLAN Benchmark, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gw Reconstruction EvalA

Evaluates the fidelity of model-agnostic Bayesian waveform reconstruction for unbound binary black hole fly-bys across different detector noise environments and frame function parameterizations. It also quantifies the astrophysical detection sensitivity and expected event rates for current and next-generation interferometers. Use when the user wants to benchmark on Simulated Hyperbolic BBH Encounters, or asks about evaluating this task. Reports overlap.

researchpythongo
0
3
Gw Glitch Mitigation EvalA

Evaluates the ability of a joint signal-glitch model to accurately recover compact binary coalescence parameters and reconstruct glitch waveforms when they overlap in real gravitational-wave detector data. It compares a baseline model ignoring glitches against a full model that jointly fits both components. Use when the user wants to benchmark on LIGO O3 data segments, or asks about evaluating this task. Reports mismatch.

researchpythongo
0
3
Guiodyssey EvalA

Evaluates multimodal agents' ability to perform cross-app GUI navigation on mobile devices by predicting correct UI actions based on screen states and task instructions. It probes spatial reasoning, action planning, and the model's capacity to leverage historical context across multiple applications. Use when the user wants to benchmark on GUIOdyssey, or asks about evaluating this task. Reports Action Matching Score (AMS).

researchpythongo
0
3
Guide Research Idea EvalA

This evaluation probes a system's ability to act as a scientific advisor by predicting whether research hypotheses will be accepted at a top-tier AI conference. It measures alignment with expert peer-review decisions using ranking-based precision and recall metrics on a held-out set of conference submissions. Use when the user wants to benchmark on ICLR 2025 Submissions Test Set, or asks about evaluating this task. Reports Top-30% Precision.

researchpythongo
0
3
Guicourse Gui Nav EvalA

Evaluates vision-language models' ability to navigate graphical user interfaces by predicting correct action types and precise screen coordinates. It probes OCR, pixel-level grounding, and multi-step task planning across web and mobile environments. Use when the user wants to benchmark on GUIAct, Mind2Web, AITW, or asks about evaluating this task. Reports StepSR.

researchpythongo
0
3
Gui Grounding Navigation EvalA

This evaluation probes a GUI agent's ability to localize UI elements via grounding and execute multi-step navigation tasks across mobile, web, and desktop platforms. It measures spatial perception, action planning consistency, and cross-platform generalization under both offline and online interaction settings. Use when the user wants to benchmark on ScreenSpot-V2, ScreenSpot-Pro, AndroidControl, AndroidWorld, ChiM-Nav, Ubu-Nav, or asks about evaluating this task. Reports success rate, Step S...

researchpythongo
0
3
Gui Grounding EvalA

Evaluates a model's ability to locate specific UI elements on screenshots based on natural language instructions. It measures both the recall of candidate generation and the precision of visual discrimination to select the correct bounding box. Use when the user wants to benchmark on MMBench-GUI, ScreenSpot-Pro, UI-Vision, ScreenSpot-v2, UI-I2E-Bench, OSWorld-G, or asks about evaluating this task. Reports Top-1 Accuracy.

researchpythongo
0
3
Gui Grounding Agent EvalA

Evaluates a GUI agent's ability to precisely locate UI elements (grounding) and execute multi-step tasks in real-world desktop/web environments. It probes spatial reasoning, text/icon matching, and long-horizon planning robustness without lookahead. Use when the user wants to benchmark on ScreenSpot-Pro, ScreenSpot-V2, OSWorld-G, OSWorld, WindowsAgentArena, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gui Ceval EvalA

Evaluates multimodal large language models and agents on Chinese mobile GUI interaction tasks. It probes atomic capabilities like visual perception, grounding, and planning, as well as end-to-end execution reliability in both offline simulation and real-device online environments. Use when the user wants to benchmark on GUI-CEval, or asks about evaluating this task. Reports Online Agent success rate.

researchpython
0
3
Gui Agent Kv EvalA

Evaluates the accuracy and efficiency of GUI agents under varying KV cache compression budgets across visual grounding, offline action prediction, and online task completion benchmarks. Use when the user wants to benchmark on ScreenSpotV2, ScreenSpot-Pro, AndroidControl, Multimodal-Mind2Web, AgentNetBench, OSWorld-Verified, or asks about evaluating this task. Reports step accuracy.

researchpythongo
0
3
Gui Agent Halluc EvalA

This protocol evaluates GUI agents on visual grounding, action execution, and hallucination rates across mobile, desktop, and web interfaces. It measures how well models localize UI elements, execute multi-step tasks under varying instruction granularities, and avoid perception or reasoning errors. Use when the user wants to benchmark on ScreenSpot-V2, ScreenSpot-Pro, AndroidControl, GUI-Odyssey, or asks about evaluating this task. Reports Action Type Accuracy (Type), Grounding Accuracy (GR),...

researchpythongo
0
3
Guardreasoner Vl EvalA

Evaluates the ability of vision-language models to detect harmful content in user prompts and AI responses across text, image, and multimodal inputs. It probes safety alignment and reasoning capabilities by measuring classification accuracy on diverse safety benchmarks. Use when the user wants to benchmark on ToxicChat, HarmBench, OpenAIModeration, AegisSafetyTest, SimpleSafetyTests, WildGuardTest, HarmImageTest, SPA-VL-Eval, SafeRLHF, BeaverTails, XSTestResponse, or asks about evaluating thi...

researchpythongo
0
3