Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,609–7,632 of 20,843 skills
Evaluates an unsupervised anomaly detection model's ability to identify potentially habitable exoplanets using planetary parameter features, deliberately excluding direct habitability labels to simulate real-world discovery scenarios. Use when the user wants to benchmark on PHL-EC (Planetary Habitability Laboratory Exoplanet Catalog), or asks about evaluating this task. Reports PHL-HEC intersection.
Evaluates high-contrast imaging algorithms' ability to detect injected exoplanet companions in real and simulated ADI sequences. It measures detection sensitivity and specificity across different noise regimes and instrument datasets, comparing performance against standard PCA and deep learning baselines. Use when the user wants to benchmark on EIDC (Exoplanet Imaging Data Challenge) Phase 1, or asks about evaluating this task. Reports F1-score.
Evaluates a cosmological galaxy simulation framework by comparing its synthetic exoplanet population demographics against real observational catalogs from the NASA Exoplanet Archive and Kepler mission. Use when the user wants to benchmark on NASA Exoplanet Archive, Kepler observations, or asks about evaluating this task. Reports planet type fraction.
Evaluates the ability of autoencoder-derived latent representations combined with various anomaly detection algorithms to identify chemically anomalous exoplanet transit spectra under realistic observational noise levels. It benchmarks reconstruction loss, 1-class SVM, K-means, and LOF across raw spectral and latent feature spaces. Use when the user wants to benchmark on Exoplanet transit spectra database, or asks about evaluating this task. Reports AUC.
This benchmark evaluates computer-use agents' ability to correctly judge whether a GUI interaction trajectory succeeds or fails, and to precisely localize the temporal window where the first error occurs. It probes spatiotemporal reasoning, visual redundancy handling, and fine-grained temporal attribution in long video trajectories. Use when the user wants to benchmark on ExeVR-Bench, or asks about evaluating this task. Reports accuracy.
Measures the wall-clock execution time of four representative Fully Homomorphic Encryption (CKKS) workloads—bootstrapping, logistic regression training, RNN inference, and ResNet-20 inference—across different GPU architectures to evaluate library performance and memory constraints. Use when the user has predictions and gold and needs to compute execution_time.
Evaluates repository-level code completion capabilities of LLMs across multiple granularities (span, line, expression, statement, function) using executable validation and string similarity metrics. Use when the user wants to benchmark on ExecRepoBench, or asks about evaluating this task. Reports Pass@1.
Evaluates foundation models on detecting, forecasting, and monitoring extreme Earth events across seven categories including heatwaves, storms, floods, and wildfires. It probes model generalizability and transferability under data scarcity, distribution shift, and severe class imbalance across heterogeneous geospatial and meteorological modalities. Use when the user wants to benchmark on ExEBench, or asks about evaluating this task. Reports Accuracy (ACC).
This benchmark evaluates a model's ability to jointly perform Chinese grammatical error correction and generate edit-wise explanations. It probes the model's capacity to identify specific error spans, assign severity levels, and provide natural-language justifications with evidence words, linguistic rules, and revision advice. Use when the user wants to benchmark on EXCGEC, or asks about evaluating this task. Reports CLEME F0.5.
Evaluates a model's ability to recognize and segment specific chart components (e.g., bars, lines, pie slices, legends, axis titles) within chart images using instance segmentation. Use when the user wants to benchmark on ExcelChart400K, or asks about evaluating this task. Reports mAP.
Evaluates explainable anomaly detection (AD) and explanation discovery (ED) capabilities on high-dimensional multivariate time series. It probes an algorithm's ability to detect range-based anomalies across four progressive difficulty levels and assesses the quality of generated explanations based on conciseness, consistency, and predictive accuracy. Use when the user wants to benchmark on Exathlon, or asks about evaluating this task. Reports Range-based Precision.
Evaluates multilingual and cross-lingual question answering capabilities on high school-level exams across multiple subjects and languages. Probes domain-specific reasoning, knowledge retrieval, and zero-shot transfer between languages with varying linguistic and subject overlaps. Use when the user wants to benchmark on EXAMS, or asks about evaluating this task. Reports accuracy.
Evaluates video-language models' ability to answer questions about expert-level physical skilled activities in long-form videos. It probes fine-grained action recognition, temporal reasoning, and domain-specific generalization across sports, bike repair, cooking, health, music, and dance. Use when the user wants to benchmark on ExAct, or asks about evaluating this task. Reports accuracy.
EWMBench evaluates embodied world models on their ability to generate videos that maintain static scene consistency, follow physically plausible motion trajectories, and align semantically with text instructions. It probes whether video generation models can produce action-consistent, task-grounded behaviors for robotic manipulation rather than just visually plausible but static or semantically drifting clips. Use when the user wants to benchmark on Agibot-World, or asks about evaluating this...
Evaluates lifelong language models' ability to update outdated world knowledge while retaining new information and avoiding catastrophic forgetting. It specifically probes temporal adaptation, numerical reasoning, and the model's capacity to forget obsolete facts during continual pretraining. Use when the user wants to benchmark on EvolvingQA, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates dynamic non-IID transfer learning on graphs by measuring how well a model adapts node classification knowledge from a source temporal graph to a target temporal graph with limited labeled samples. It probes the model's ability to handle evolving graph structures, domain discrepancies, and temporal dependencies across heterogeneous datasets. Use when the user wants to benchmark on DBLP-3, DBLP-5, HCP, or asks about evaluating this task. Reports AUC.
Evaluates the ability of Evolutive RNNs (EvoRNNs) to model long-range dependencies in sequential data. It compares EvoRNNs against standard RNN baselines on language modeling and sequential recommendation tasks, emphasizing both predictive accuracy and computational efficiency. Use when the user wants to benchmark on LM1B (Billion word dataset), Sequential recommendation dataset, or asks about evaluating this task. Reports MAP@20.
Evaluates the ability of neural radiance field (NeRF) models to accurately reconstruct 3D scenes from 2D images while simultaneously quantifying both aleatoric (data noise) and epistemic (model ignorance) uncertainties. It probes whether uncertainty estimates reliably correlate with actual rendering errors and calibration across varying scene conditions and data sparsity. Use when the user wants to benchmark on Light Field (LF), Local Light Field Fusion (LLFF), RobustNeRF, or asks about evalu...
This evaluation probes the fidelity of an LLM-assisted pipeline in extracting structured, PICO-style evidence nodes from unstructured full-text biomedical literature, and the precision of subsequent entity normalization against a reference resource. Use when the user wants to benchmark on HCC and CRC PubMed corpus, or asks about evaluating this task. Reports field-level extraction accuracy.
This benchmark evaluates a model's ability to identify and classify clinical evidence spans within randomized controlled trial (RCT) documents. Specifically, it probes whether a given evidence span supports a significantly decreased, no significant difference, or significantly increased outcome relative to a clinical intervention. Use when the user wants to benchmark on Evidence Inference, or asks about evaluating this task. Reports macro-averaged F1.
EveTAR evaluates Arabic information retrieval systems across four tasks: event detection, ad-hoc search, tweet timeline generation, and real-time summarization. It probes a model's ability to retrieve, cluster, and summarize relevant Arabic tweets in response to event-driven queries or topics. Use when the user wants to benchmark on EveTAR, or asks about evaluating this task. Reports recall.
Evaluates large language models' ability to infer natural language event sequences from real-valued time series data (specifically win probabilities in sports). It probes causal reasoning, temporal context understanding, and the model's capacity to distinguish underlying time series dynamics from linguistic descriptions. Use when the user wants to benchmark on NBA & NFL Event Inference Benchmark, or asks about evaluating this task. Reports accuracy.
Evaluates an LLM's ability to perform spatial and contextual reasoning for multi-agent planning in 3D scenes. It tests object arrangement, regional context alignment, and scene state tracking to generate plausible character actions and positions. Use when the user wants to benchmark on Event-Driven Storytelling Benchmark, or asks about evaluating this task. Reports success rate.
Evaluates vision-language models' ability to recognize evoked emotions from images in a zero-shot setting. It probes their robustness to prompt perturbations and measures sentiment bias in predicting positive vs. negative emotions. Use when the user wants to benchmark on EvE, or asks about evaluating this task. Reports weighted F1 score.