Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,843
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,609–7,632 of 20,843 skills

Exoplanet Habitability Anomaly EvalA

Evaluates an unsupervised anomaly detection model's ability to identify potentially habitable exoplanets using planetary parameter features, deliberately excluding direct habitability labels to simulate real-world discovery scenarios. Use when the user wants to benchmark on PHL-EC (Planetary Habitability Laboratory Exoplanet Catalog), or asks about evaluating this task. Reports PHL-HEC intersection.

researchpythongo
0
3
Exoplanet Detection EvalA

Evaluates high-contrast imaging algorithms' ability to detect injected exoplanet companions in real and simulated ADI sequences. It measures detection sensitivity and specificity across different noise regimes and instrument datasets, comparing performance against standard PCA and deep learning baselines. Use when the user wants to benchmark on EIDC (Exoplanet Imaging Data Challenge) Phase 1, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Exoplanet Demographics EvalA

Evaluates a cosmological galaxy simulation framework by comparing its synthetic exoplanet population demographics against real observational catalogs from the NASA Exoplanet Archive and Kepler mission. Use when the user wants to benchmark on NASA Exoplanet Archive, Kepler observations, or asks about evaluating this task. Reports planet type fraction.

researchpythongo
0
3
Exoplanet Anomaly Detection EvalA

Evaluates the ability of autoencoder-derived latent representations combined with various anomaly detection algorithms to identify chemically anomalous exoplanet transit spectra under realistic observational noise levels. It benchmarks reconstruction loss, 1-class SVM, K-means, and LOF across raw spectral and latent feature spaces. Use when the user wants to benchmark on Exoplanet transit spectra database, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Exevr Bench EvalA

This benchmark evaluates computer-use agents' ability to correctly judge whether a GUI interaction trajectory succeeds or fails, and to precisely localize the temporal window where the first error occurs. It probes spatiotemporal reasoning, visual redundancy handling, and fine-grained temporal attribution in long video trajectories. Use when the user wants to benchmark on ExeVR-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Execution TimeA

Measures the wall-clock execution time of four representative Fully Homomorphic Encryption (CKKS) workloads—bootstrapping, logistic regression training, RNN inference, and ResNet-20 inference—across different GPU architectures to evaluate library performance and memory constraints. Use when the user has predictions and gold and needs to compute execution_time.

researchpythongo
0
3
Execrepo Bench EvalA

Evaluates repository-level code completion capabilities of LLMs across multiple granularities (span, line, expression, statement, function) using executable validation and string similarity metrics. Use when the user wants to benchmark on ExecRepoBench, or asks about evaluating this task. Reports Pass@1.

researchpythongo
0
3
Exebench EvalA

Evaluates foundation models on detecting, forecasting, and monitoring extreme Earth events across seven categories including heatwaves, storms, floods, and wildfires. It probes model generalizability and transferability under data scarcity, distribution shift, and severe class imbalance across heterogeneous geospatial and meteorological modalities. Use when the user wants to benchmark on ExEBench, or asks about evaluating this task. Reports Accuracy (ACC).

researchpythongo
0
3
Excgex EvalA

This benchmark evaluates a model's ability to jointly perform Chinese grammatical error correction and generate edit-wise explanations. It probes the model's capacity to identify specific error spans, assign severity levels, and provide natural-language justifications with evidence words, linguistic rules, and revision advice. Use when the user wants to benchmark on EXCGEC, or asks about evaluating this task. Reports CLEME F0.5.

researchpythongo
0
3
Excelchart400k EvalA

Evaluates a model's ability to recognize and segment specific chart components (e.g., bars, lines, pie slices, legends, axis titles) within chart images using instance segmentation. Use when the user wants to benchmark on ExcelChart400K, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
Exathlon EvalA

Evaluates explainable anomaly detection (AD) and explanation discovery (ED) capabilities on high-dimensional multivariate time series. It probes an algorithm's ability to detect range-based anomalies across four progressive difficulty levels and assesses the quality of generated explanations based on conciseness, consistency, and predictive accuracy. Use when the user wants to benchmark on Exathlon, or asks about evaluating this task. Reports Range-based Precision.

researchpythongo
0
3
Exams Qa EvalA

Evaluates multilingual and cross-lingual question answering capabilities on high school-level exams across multiple subjects and languages. Probes domain-specific reasoning, knowledge retrieval, and zero-shot transfer between languages with varying linguistic and subject overlaps. Use when the user wants to benchmark on EXAMS, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Exact EvalA

Evaluates video-language models' ability to answer questions about expert-level physical skilled activities in long-form videos. It probes fine-grained action recognition, temporal reasoning, and domain-specific generalization across sports, bike repair, cooking, health, music, and dance. Use when the user wants to benchmark on ExAct, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ewmbench EvalA

EWMBench evaluates embodied world models on their ability to generate videos that maintain static scene consistency, follow physically plausible motion trajectories, and align semantically with text instructions. It probes whether video generation models can produce action-consistent, task-grounded behaviors for robotic manipulation rather than just visually plausible but static or semantically drifting clips. Use when the user wants to benchmark on Agibot-World, or asks about evaluating this...

researchpythongit
0
3
Evolvingqa EvalA

Evaluates lifelong language models' ability to update outdated world knowledge while retaining new information and avoiding catastrophic forgetting. It specifically probes temporal adaptation, numerical reasoning, and the model's capacity to forget obsolete facts during continual pretraining. Use when the user wants to benchmark on EvolvingQA, or asks about evaluating this task. Reports Exact Match (EM).

researchpythongo
0
3
Evoluenet Dynamic Graph Transfer EvalA

Evaluates dynamic non-IID transfer learning on graphs by measuring how well a model adapts node classification knowledge from a source temporal graph to a target temporal graph with limited labeled samples. It probes the model's ability to handle evolving graph structures, domain discrepancies, and temporal dependencies across heterogeneous datasets. Use when the user wants to benchmark on DBLP-3, DBLP-5, HCP, or asks about evaluating this task. Reports AUC.

researchpythonnode
0
3
Evo Rnn EvalA

Evaluates the ability of Evolutive RNNs (EvoRNNs) to model long-range dependencies in sequential data. It compares EvoRNNs against standard RNN baselines on language modeling and sequential recommendation tasks, emphasizing both predictive accuracy and computational efficiency. Use when the user wants to benchmark on LM1B (Billion word dataset), Sequential recommendation dataset, or asks about evaluating this task. Reports MAP@20.

researchpythongo
0
3
Evidential Nerf EvalA

Evaluates the ability of neural radiance field (NeRF) models to accurately reconstruct 3D scenes from 2D images while simultaneously quantifying both aleatoric (data noise) and epistemic (model ignorance) uncertainties. It probes whether uncertainty estimates reliably correlate with actual rendering errors and calibration across varying scene conditions and data sparsity. Use when the user wants to benchmark on Light Field (LF), Local Light Field Fusion (LLFF), RobustNeRF, or asks about evalu...

researchpython
0
3
Evidencenet Extraction EvalA

This evaluation probes the fidelity of an LLM-assisted pipeline in extracting structured, PICO-style evidence nodes from unstructured full-text biomedical literature, and the precision of subsequent entity normalization against a reference resource. Use when the user wants to benchmark on HCC and CRC PubMed corpus, or asks about evaluating this task. Reports field-level extraction accuracy.

researchpythongo
0
3
Evidence Inference EvalA

This benchmark evaluates a model's ability to identify and classify clinical evidence spans within randomized controlled trial (RCT) documents. Specifically, it probes whether a given evidence span supports a significantly decreased, no significant difference, or significantly increased outcome relative to a clinical intervention. Use when the user wants to benchmark on Evidence Inference, or asks about evaluating this task. Reports macro-averaged F1.

researchpythongo
0
3
Evetar EvalA

EveTAR evaluates Arabic information retrieval systems across four tasks: event detection, ad-hoc search, tweet timeline generation, and real-time summarization. It probes a model's ability to retrieve, cluster, and summarize relevant Arabic tweets in response to event-driven queries or topics. Use when the user wants to benchmark on EveTAR, or asks about evaluating this task. Reports recall.

researchpythongo
0
3
Event Inference EvalA

Evaluates large language models' ability to infer natural language event sequences from real-valued time series data (specifically win probabilities in sports). It probes causal reasoning, temporal context understanding, and the model's capacity to distinguish underlying time series dynamics from linguistic descriptions. Use when the user wants to benchmark on NBA & NFL Event Inference Benchmark, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Event Driven Storytelling EvalA

Evaluates an LLM's ability to perform spatial and contextual reasoning for multi-agent planning in 3D scenes. It tests object arrangement, regional context alignment, and scene state tracking to generate plausible character actions and positions. Use when the user wants to benchmark on Event-Driven Storytelling Benchmark, or asks about evaluating this task. Reports success rate.

researchpython
0
3
Eve Emotion Recognition EvalA

Evaluates vision-language models' ability to recognize evoked emotions from images in a zero-shot setting. It probes their robustness to prompt perturbations and measures sentiment bias in predicting positive vs. negative emotions. Use when the user wants to benchmark on EvE, or asks about evaluating this task. Reports weighted F1 score.

researchpythongo
0
3