All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,385
- 892
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,889–9,912 of 21,385 skills
- Abraham EvalEvaluates molecular generative models by assessing their ability to recreate known ligands, predict drug-target affinity, and bind to target proteins via molecular docking. It probes the biological relevance and structural fidelity of de novo generated molecules across multiple protein targets. Use when the user wants to benchmark on ABRAHAM, or asks about evaluating this task. Reports ROOM recreation metric.Votes: 0GitHub stars: 3
- Abnormal Driving Detection EvalDetects abnormal driving behaviors in naturalistic driving data using event-level safety indicators and motion features. It evaluates a semi-supervised machine learning model's ability to distinguish between normal and anomalous driving events based on vehicle dynamics and temporal proximity metrics. Use when the user wants to benchmark on Naturalistic Driving Dataset, or asks about evaluating this task. Reports F1-score.Votes: 0GitHub stars: 3
- Abdomenct1k EvalThis benchmark evaluates the ability of 3D medical image segmentation models to accurately delineate abdominal organs (liver, kidney, spleen, pancreas) under clinically challenging conditions. It specifically probes generalization across unseen medical centers, CT contrast phases, and severe pathologies like tumors, while measuring both volumetric overlap and boundary precision. Use when the user wants to benchmark on AbdomenCT-1K, or asks about evaluating this task. Reports DSC.Votes: 0GitHub stars: 3
- Abcfair EvalEvaluates the trade-off between predictive performance and fairness across diverse real-world settings. It probes how different intervention stages, sensitive feature compositions, fairness notions, and output distributions impact a model's ability to satisfy fairness constraints while maintaining accuracy. Use when the user wants to benchmark on SchoolPerformance, ACSPublicCoverage, or asks about evaluating this task. Reports AUROC.Votes: 0GitHub stars: 3
- Abc EvalThis benchmark evaluates large language models' ability to understand symbolic music and follow instructions using text-based ABC notation. It probes capabilities ranging from basic syntax parsing and error detection to segment-level reasoning and sequence-level musical analysis like genre or emotion recognition. Use when the user wants to benchmark on ABC-Eval, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Aa Omniscience EvalEvaluates large language models' factual recall and knowledge calibration across domain-specific questions. It measures how reliably models provide correct answers versus hallucinating or abstaining when uncertain, highlighting the gap between raw accuracy and factual reliability. Use when the user wants to benchmark on AA-Omniscience, or asks about evaluating this task. Reports Omniscience Index.Votes: 0GitHub stars: 3
- A3 EvalEvaluates mobile GUI agents on completing multi-step tasks across 20 real-world Android applications. It probes both final task completion capability and the agent's ability to navigate intermediate essential states without getting stuck or making terminal errors. Use when the user wants to benchmark on A3, or asks about evaluating this task. Reports Task Success Rate (SR).Votes: 0GitHub stars: 3
- A2seek EvalEvaluates multimodal models' ability to detect, localize, and semantically reason about anomalies in aerial drone-view videos. It probes spatial grounding accuracy, temporal anomaly detection, and the generation of contextually grounded natural language explanations. Use when the user wants to benchmark on A2Seek, or asks about evaluating this task. Reports AP_c, mIoU.Votes: 0GitHub stars: 3
- TTSDSMeasures the distributional distance between synthetic and real speech across five key factors: environment, speaker identity, prosody, intelligibility, and general speech distribution. It evaluates TTS system quality without relying on subjective Mean Opinion Scores (MOS) or simple mean-based metrics. Use when the user has predictions and gold and needs to compute TTSDS.Votes: 0GitHub stars: 3
- TECEvaluates the joint optimization of AI service placement and resource allocation in mobile edge computing by measuring the trade-off between computation time and energy consumption across varying network scales and task characteristics. Use when the user has predictions and gold and needs to compute TEC.Votes: 0GitHub stars: 3
- TCTBEvaluates the throughput and resource allocation efficiency of RIS-aided mobile edge computing systems by measuring the total computation task bits successfully completed under varying network conditions. Use when the user has predictions and gold and needs to compute TCTB.Votes: 0GitHub stars: 3
- SDREvaluates how effectively audio separation models isolate specific stems from a mixed recording while preserving signal integrity. It quantifies the ratio of target source energy to residual interference and distortion energy. Use when the user has predictions and gold and needs to compute SDR.Votes: 0GitHub stars: 3
- PSNREvaluates the trade-off between file size reduction and image fidelity when encoding radio astronomy data using JPEG2000. It benchmarks both lossless and lossy compression modes to determine the compression ratio at which visual artifacts first appear. Use when the user has predictions and gold and needs to compute PSNR.Votes: 0GitHub stars: 3
- NDCG@10Evaluates how well internal model representations (hidden states) can predict token-level information importance in summarization tasks. It probes whether specific transformer layers or cross-layer combinations encode salience distributions consistent with empirical importance derived from summary persistence. Use when the user has predictions and gold and needs to compute NDCG@10.Votes: 0GitHub stars: 3
- MENLIEvaluates the robustness and alignment with human judgment of reference-based and reference-free evaluation metrics for machine translation and summarization, particularly under adversarial conditions. Use when the user has predictions and gold and needs to compute Pearson correlation.Votes: 0GitHub stars: 3
- EASThis evaluation validates the Emotional Attitude Score (EAS) metric by measuring its consistency with human judgment on word-level sentiment polarity. It probes whether the metric's pseudo-log-likelihood-based scores reliably capture positive or negative emotional attitudes toward ambiguous attitude words in gender-inclusive contexts. Use when the user has predictions and gold and needs to compute EAS.Votes: 0GitHub stars: 3
- APESsrcEvaluates the faithfulness of abstractive summaries by verifying if factual claims (masked as cloze questions) in the reference summary can be correctly answered using only the generated summary, compared against a gold-standard answer derived from the source context. Use when the user has predictions and gold and needs to compute APESsrc.Votes: 0GitHub stars: 3
- 6dof Visual Localization EvalEvaluates a model's ability to estimate 6-degree-of-freedom camera poses for query images against a reference 3D model, specifically testing robustness to drastic changes in lighting (day/night), weather, and seasonal vegetation. Use when the user wants to benchmark on Aachen Day-Night, RobotCar Seasons, CMU Seasons, or asks about evaluating this task. Reports translation error and rotation error.Votes: 0GitHub stars: 3
- 6dof Camera Tracking EvalEvaluates the tracking accuracy of a 6-DoF autonomous camera algorithm in a simulated surgical environment. It also measures how different camera control strategies impact human rater accuracy when assessing surgical skill from video. Use when the user wants to benchmark on da Vinci wire chaser simulation, or asks about evaluating this task. Reports assessment_error.Votes: 0GitHub stars: 3
- 5g Madrl Sumrate EvalEvaluates the ability of a multi-agent deep reinforcement learning framework to optimize the 3D placement and trajectory of mobile access points in dynamic 5G networks, balancing sum-rate maximization against user mobility and interference. Use when the user wants to benchmark on Custom 5G Network Simulation, or asks about evaluating this task. Reports sum-rate.Votes: 0GitHub stars: 3
- 4seasons EvalEvaluates visual SLAM and long-term localization for autonomous driving under challenging cross-season, multi-weather, and long-term environmental changes. Specifically probes visual odometry, global place recognition, and map-based visual localization capabilities. Use when the user wants to benchmark on 4Seasons, or asks about evaluating this task. Reports horizontal RMSE.Votes: 0GitHub stars: 3
- 3mdbench EvalEvaluates Large Vision-Language Models in realistic telemedicine consultations by simulating multi-agent dialogues between a doctor and a temperament-based patient. It probes diagnostic accuracy from multimodal inputs (images + text) and assesses clinical competence and dialogue quality. Use when the user wants to benchmark on 3MDBench, or asks about evaluating this task. Reports F1 Score.Votes: 0GitHub stars: 3
- 3dses EvalSemantic segmentation of indoor Terrestrial Laser Scanning (TLS) point clouds. It probes a model's ability to classify 3D points into semantic categories (e.g., furniture, structural elements, clutter) using geometric coordinates and optionally Lidar intensity features. Use when the user wants to benchmark on 3DSES, or asks about evaluating this task. Reports mIoU.Votes: 0GitHub stars: 3
- 3doc Bench EvalThis benchmark evaluates a model's ability to generate text-to-image outputs that strictly adhere to 3D layout constraints, handle complex inter-object occlusions, and maintain correct object orientations and visibility orders. It probes depth-consistent scene composition, attribute binding to specific objects, and overall image fidelity under varying camera viewpoints. Use when the user wants to benchmark on 3DOc-Bench, or asks about evaluating this task. Reports depth ordering.Votes: 0GitHub stars: 3