Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,849
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,753–7,776 of 20,849 skills

Eicu Crd Clinical Bench EvalA

Evaluates machine learning models on four critical care prediction tasks using the multi-centre eICU-CRD dataset: in-hospital mortality, remaining length of stay, patient phenotyping, and physiologic decompensation. It probes the models' ability to handle longitudinal clinical data, compare categorical vs numerical feature representations, and generalize across multi-centre settings. Use when the user wants to benchmark on eICU-CRD, or asks about evaluating this task. Reports AUROC.

researchpythongo
0
3
Eicap Bench EvalA

Evaluates large language models' emotional intelligence (EI) capabilities across a four-layer taxonomy: emotional tracking, cause inference, appraisal, and emotionally appropriate response generation. It probes fine-grained subcategories including cultural sensitivity, valence judgment, and uncertainty calibration using multi-turn conversational contexts. Use when the user wants to benchmark on EICap-Bench, or asks about evaluating this task. Reports macro-average accuracy.

researchpythongo
0
3
Ehrscl 2024 EvalA

Evaluates a system's ability to translate natural language clinical questions into executable SQL and retrieve accurate results from a specialized electronic health record database. It probes complex temporal reasoning, clinical constraint handling, and semantic equivalence in text-to-SQL generation. Use when the user wants to benchmark on EHRSQL 2024, or asks about evaluating this task. Reports execution accuracy.

researchpythonsql
0
3
Ehrr1 EvalA

Evaluates a language model's ability to perform clinical decision-making and risk prediction using longitudinal electronic health record (EHR) data. It probes the model's capacity for multi-label entity recommendation, binary outcome forecasting, and generalization across different healthcare systems and diagnostic granularities. Use when the user wants to benchmark on EHR-Bench, MIMIC-IV-CDM, EHRSHOT, or asks about evaluating this task. Reports F1 score, AUROC.

researchpythongo
0
3
Ehrnoteqa EvalA

Evaluates large language models' ability to perform patient-specific clinical reasoning by synthesizing information from multiple electronic health record (EHR) discharge summaries to answer medical questions. It specifically tests multi-document clinical analysis and automated medical model evaluation using structured multi-choice or free-text formats. Use when the user wants to benchmark on EHRNoteQA, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Ehr Clinical Outcome Prediction EvalA

This benchmark evaluates clinical outcome prediction models across three distinct EHR data representations (multivariate time-series, event streams, and textual event streams). It probes how well different architectures handle sparse, irregular longitudinal patient data and varying feature missingness rates in both acute ICU and long-term care settings. Use when the user wants to benchmark on MIMIC-IV, EHRSHOT, or asks about evaluating this task. Reports F1 score, AUROC, AUPRC.

researchpythongit
0
3
Egtr Sgg EvalA

Evaluates a model's ability to detect objects and predict relational triplets (subject-predicate-object) in natural images. It probes both object detection accuracy and scene graph generation quality under graph constraints and standard recall/mAP metrics. Use when the user wants to benchmark on Visual Genome, Open Image V6, or asks about evaluating this task. Reports Recall@k (R@k), micro-R@50.

researchpythongo
0
3
Egoxtreme EvalA

Evaluates the robustness of 6D object pose estimation models under extreme real-world visual conditions, including severe motion blur, dynamic lighting, and smoke. It also benchmarks temporal tracking strategies in highly dynamic egocentric scenarios to assess motion-aware inference capabilities. Use when the user wants to benchmark on EgoXtreme, or asks about evaluating this task. Reports ADD(-S) recall.

researchpythongo
0
3
Egotraj Bench EvalA

Evaluates the robustness of trajectory prediction models when historical observations are corrupted by realistic ego-view perception noise (occlusions, ID switches, ego-motion drift) compared to clean bird's-eye-view ground truth. Use when the user wants to benchmark on EgoTraj-TBD, or asks about evaluating this task. Reports minADE@K, minFDE@K.

researchpythongo
0
3
Egoscreen Emotion EvalA

Evaluates a model's ability to predict human emotional responses to movie scenes from an egocentric, first-person screen-view perspective. It probes multimodal long-context reasoning by combining visual frames, audio cues, and narrative summaries to handle domain shifts from cinematic to realistic viewing conditions. Use when the user wants to benchmark on EgoScreen-Emotion (ESE), or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Egoschema EvalA

Probes long-term visual memory and temporal reasoning in video-language models by requiring them to answer multiple-choice questions about very long-form videos. It measures the model's ability to retain and retrieve information across extended durations without relying on short clip analysis. Use when the user wants to benchmark on EgoSchema, or asks about evaluating this task. Reports QA Accuracy.

researchpythongo
0
3
Egonormia EvalA

Evaluates vision-language models' ability to understand and reason about physical-social norms in egocentric video scenarios. It probes whether models can correctly select normative actions, justify them, and identify plausible alternatives in conflict-prone situations. Use when the user wants to benchmark on EgoNormia, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Egomem EvalA

Evaluates a lifelong memory agent's ability to perform real-time audiovisual user retrieval, detect dialog session boundaries in continuous streams, and generate personalized, fact-consistent responses in full-duplex omnimodal interactions. Use when the user wants to benchmark on LFW, VoxCeleb, EgoMem Custom Text Retrieval, EgoMem Episodic Trigger, or asks about evaluating this task. Reports pass@5, Fact Score.

researchpythongo
0
3
Egohumans EvalA

Evaluates the ability of models to perform robust multi-human tracking and identity association in unconstrained egocentric 3D environments. It probes how well algorithms handle severe occlusions, dynamic activities, and camera-agnostic spatial reasoning when fusing egocentric and secondary views. Use when the user wants to benchmark on EgoHumans, or asks about evaluating this task. Reports IDF1.

researchpythongo
0
3
Egoavu Bench EvalA

Evaluates multimodal large language models' ability to perform joint audio-visual reasoning on egocentric videos, including action/object/sound recognition, temporal reasoning, hallucination detection, and dense audio-visual narration. It specifically probes whether models can correctly associate environmental sounds with their visual sources and maintain temporal alignment without relying heavily on visual cues. Use when the user wants to benchmark on EgoAVU-Bench, or asks about evaluating t...

researchpythongo
0
3
Ego3d Bench EvalA

Evaluates 3D spatial reasoning and multi-view understanding in Vision-Language Models, specifically testing ego-centric distance estimation, object localization, motion tracking, travel time estimation, and relative location reasoning across multiple camera views. Use when the user wants to benchmark on Ego3D-Bench, or asks about evaluating this task. Reports Accuracy (%), RMSE.

researchpythongo
0
3
Ego Walk Nav EvalA

Evaluates the ability of visual navigation models to predict future robot trajectories from egocentric video frames and context history. It probes scale-invariant trajectory prediction and alignment with human navigation behavior under domain shift conditions. Use when the user wants to benchmark on EgoWalk, or asks about evaluating this task. Reports MSE.

researchpythongo
0
3
Ego Instructor EvalA

This evaluation protocol assesses a retrieval-augmented egocentric video captioning framework. It probes the model's ability to perform cross-view video-text and video-video retrieval, answer multiple-choice questions based on video-text alignment, and generate accurate egocentric video captions using retrieved exocentric instructional videos as references. Use when the user wants to benchmark on EK100 MIR, EgoMCQ, SummMCQ, YouCook2-Clip, YouCook2-Video, CharadesEgo, EgoLearner-MCQ, Ego4d coo...

researchpythongo
0
3
Egmm Corpus EvalA

Probes zero-shot visual-language alignment and cultural recognition capabilities of vision-language models on Egyptian cultural concepts. It measures how well models can classify images into specific cultural categories and retrieve matching text descriptions without fine-tuning. Use when the user wants to benchmark on EgMM-Corpus, or asks about evaluating this task. Reports Acc@1.

researchpythongo
0
3
Egida Safety EvalA

Evaluates the robustness of LLMs against jailbreaking attacks after safety alignment. It measures how well models refuse harmful prompts across diverse topics and attack styles, while also tracking unintended side effects like over-refusal and general capability degradation. Use when the user wants to benchmark on Egida, or asks about evaluating this task. Reports ASR.

researchpythonperformance
0
3
Ege Math Assessment EvalA

This benchmark evaluates vision-language models' ability to assess handwritten mathematical solutions against a standardized educational rubric. It probes the models' capacity for error diagnosis, step-by-step reasoning alignment, and accurate grade assignment under varying levels of contextual guidance. Use when the user wants to benchmark on EGE-Math Solutions Assessment Benchmark, or asks about evaluating this task. Reports final_score.

researchpythongo
0
3
Efok Cqa EvalA

Evaluates knowledge graph complex query answering models on existential first-order (EFO) queries with multiple free variables and complex structures (cycles, multi-hop), testing their ability to handle combinatorially hard queries beyond simple set operations. Use when the user wants to benchmark on EFO_k-CQA, or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Efficient Bert EvalA

Evaluates the performance of efficiently trained BERT models (via Mixture-of-Supernets) on downstream natural language understanding tasks. It probes the trade-off between model size, training compute, and accuracy compared to standalone pretraining and other NAS baselines. Use when the user wants to benchmark on GLUE benchmark, or asks about evaluating this task. Reports Avg. GLUE.

researchpythonperformance
0
3
Effective DimensionalityA

Effective Dimensionality (ED) quantifies the number of independent signals or latent axes captured by a benchmark, measuring how much redundancy exists across its tasks. It probes whether a benchmark's claimed breadth actually reflects diverse evaluation dimensions or merely correlated task performance. Use when the user has predictions and gold and needs to compute Effective Dimensionality (ED).

researchpythongo
0
3