Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,853
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,777–7,800 of 20,853 skills

Efficient Bert EvalA

Evaluates the performance of efficiently trained BERT models (via Mixture-of-Supernets) on downstream natural language understanding tasks. It probes the trade-off between model size, training compute, and accuracy compared to standalone pretraining and other NAS baselines. Use when the user wants to benchmark on GLUE benchmark, or asks about evaluating this task. Reports Avg. GLUE.

researchpythonperformance
0
3
Effective DimensionalityA

Effective Dimensionality (ED) quantifies the number of independent signals or latent axes captured by a benchmark, measuring how much redundancy exists across its tasks. It probes whether a benchmark's claimed breadth actually reflects diverse evaluation dimensions or merely correlated task performance. Use when the user has predictions and gold and needs to compute Effective Dimensionality (ED).

researchpythongo
0
3
Ef4inca Nowcast EvalA

Evaluates the capability of spatiotemporal Transformer models to nowcast convective precipitation up to 90 minutes ahead using multi-source meteorological data. It probes the model's ability to fuse satellite infrared, radar, and NWP inputs to accurately predict the initiation, location, and intensity of rapidly evolving convective cells. Use when the user wants to benchmark on Austria convective precipitation dataset, or asks about evaluating this task. Reports Critical Success Index (CSI).

researchpythongit
0
3
Eeg Ssl Emotion EvalA

Evaluates semi-supervised EEG-based emotion recognition under extreme label scarcity. It probes the model's ability to leverage unlabeled data via representation alignment while maintaining classification performance across subject-dependent and subject-independent protocols. Use when the user wants to benchmark on SEED, SEED-IV, SEED-V, AMIGOS, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Eeg Asr Noisy Speech EvalA

Evaluates end-to-end continuous speech recognition models using only electroencephalography (EEG) signals, and assesses robustness to background noise by fusing EEG with acoustic features. Use when the user wants to benchmark on Database A, Database B, or asks about evaluating this task. Reports Word Error Rate (WER).

researchpythondatabase
0
3
Eefsuva EvalA

Evaluates LLMs' ability to solve nonstandard mathematical Olympiad problems from Eastern European and former Soviet Union competitions. It probes genuine mathematical reasoning and adaptability by testing whether models can solve problems from first principles rather than relying on cached solutions or pattern matching from familiar Western benchmarks. Use when the user wants to benchmark on EEFSUVA, or asks about evaluating this task. Reports pass rate.

researchpythongo
0
3
Eee Bench EvalA

Probes multimodal reasoning and visual diagram interpretation in electrical and electronics engineering. It tests whether models can integrate complex circuit and system diagrams with textual problem descriptions to apply domain-specific knowledge and perform accurate calculations or logical deductions. Use when the user wants to benchmark on EEE-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ee Power Control EvalA

This evaluation probes the energy efficiency and feasibility of power control algorithms in 5G massive MIMO and relay-assisted interference networks. It measures how well centralized and distributed algorithms maximize Global Energy Efficiency (GEE) while satisfying minimum per-user rate constraints under hardware impairments and Rayleigh fading. Use when the user wants to benchmark on Hardware-Impaired Massive MIMO System, Relay-assisted OFDMA interference network, or asks about evaluating t...

researchpythongo
0
3
Editverse Bench EvalA

Evaluates instruction-based video editing capabilities, including text alignment, temporal consistency, and editing faithfulness across diverse resolutions and orientations. It probes the model's ability to follow complex editing prompts while preserving unedited regions and maintaining high video quality. Use when the user wants to benchmark on EditVerseBench, or asks about evaluating this task. Reports VLM evaluation (Editing Quality).

researchpythongo
0
3
Editreward EvalA

Evaluates the quality and human alignment of instruction-guided image editing models. It measures how well generated images match user instructions and visual realism, as well as how accurately models rank pairs of edited images according to human preferences. Use when the user wants to benchmark on ImagenHub, GenAI-Bench, AURORA-Bench, EditReward-Bench, GEdit-Bench, or asks about evaluating this task. Reports Spearman rank correlation, Pair-wise prediction accuracy.

researchpythongo
0
3
Editeval EvalA

Evaluates instruction-based text editing capabilities across modular tasks such as simplification, fluency, coherence, paraphrasing, neutralization, and information updating. It measures how well models follow prompts to improve or modify text while preserving intended meaning. Use when the user wants to benchmark on EditEval Benchmark, or asks about evaluating this task. Reports SARI.

researchpythongit
0
3
Ediref Erc EfrA

Evaluates emotion recognition and emotion-flip reasoning in multi-party conversations, specifically identifying trigger utterances that cause emotional shifts in both code-mixed (Hindi-English) and monolingual English dialogues. Use when the user wants to benchmark on E-MaSaC, MELD-FR, or asks about evaluating this task. Reports weighted F1, F1 score for trigger utterances.

researchpythongit
0
3
Edge Llm Inference EvalA

Evaluates the feasibility and performance of deploying various LLMs on CPU-only edge hardware (Raspberry Pi 5 clusters) by measuring inference speed, resource consumption, and reasoning accuracy under constrained conditions. Use when the user wants to benchmark on OpenAssistant/oasst1 (subset), Winogrande, or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Edge Lm Inference EvalA

Evaluates the feasibility and performance trade-offs of running small generative language models on edge hardware. It probes memory constraints, inference latency, token throughput, and energy efficiency across different quantization schemes and system configurations. Use when the user wants to benchmark on None (system-level inference benchmark), or asks about evaluating this task. Reports generation_throughput.

researchpythongo
0
3
Edge Llm Inference BenchmarkA

Evaluates the trade-offs between token throughput, latency, energy efficiency, and physical footprint when deploying compact LLMs on various IoT-grade single-board computers with different hardware accelerators (CPU, NPU, GPU). Use when the user has predictions and gold and needs to compute Throughput (tokens/s), Time-to-first-token (TTFT), Energy per million tokens (MJ/Mtok).

researchpythongo
0
3
Edge Llm Energy Accuracy EvalA

Evaluates the trade-offs between quantization levels, model architectures, and task characteristics on energy efficiency and output accuracy for LLMs deployed on edge hardware. Use when the user wants to benchmark on bigbenchhard, commonsenseqa, gsm8k, humaneval, truthfulqa, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Edge Case Detection EvalA

Tests the model's ability to classify whether a respondent's message represents an edge case that falls outside the scope of existing coordination policies and requires user escalation. Use when the user wants to benchmark on Edge Case Detection Test Suite, or asks about evaluating this task. Reports Accuracy.

researchpythonrust
0
3
Edacc EvalA

Evaluates automatic speech recognition (ASR) models on naturalistic, conversational English speech with diverse international accents. It probes the robustness of state-of-the-art ASR systems to real-world speaking conditions and accent variation compared to read-speech benchmarks like LibriSpeech. Use when the user wants to benchmark on EdAcc, or asks about evaluating this task. Reports WER.

researchpython
0
3
Ecvr Prediction EvalA

Evaluates the ability to predict click-through, conversion, and effective conversion rates in a large-scale e-commerce recommender system. It specifically probes how well models handle cascade delayed feedback, sample selection bias, and data sparsity when predicting user purchase and refund behaviors. Use when the user wants to benchmark on Alibaba Production Dataset, or asks about evaluating this task. Reports AUC.

researchpythonperformance
0
3
Ecovnet Covidx EvalA

Evaluates the ability of deep convolutional neural networks to classify chest X-ray images into three categories: COVID-19, normal, and pneumonia. It probes robustness under class imbalance and tests the effectiveness of ensemble learning strategies (hard vs. soft voting) combined with data augmentation. Use when the user wants to benchmark on COVIDx, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ecosched Hpc Scheduling EvalA

Evaluates an online co-scheduling framework's ability to jointly optimize GPU count selection and job packing to minimize energy consumption while maintaining performance across diverse multi-GPU HPC workloads. It probes the scheduler's capacity to handle non-linear scaling, NUMA-aware placement, and dynamic workload packing without prior knowledge of exact runtimes. Use when the user wants to benchmark on Multi-GPU Benchmark Suite, or asks about evaluating this task. Reports Energy Saving.

researchpythongo
0
3
Econlogicqa EvalA

This benchmark evaluates large language models' ability to perform economic sequential reasoning by logically ordering interconnected business and supply chain events. It probes multi-event causality and temporal reasoning beyond simple chronological sorting, requiring models to understand complex economic narratives. Use when the user wants to benchmark on EconLogicQA, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Ecoli Periodicity Detection EvalA

Evaluates the ability to detect periodic outlier patterns in protein sequence time-series data. It measures the statistical significance and reliability of discovered patterns compared to a baseline algorithm. Use when the user wants to benchmark on E.Coli, or asks about evaluating this task. Reports Surprise score.

researchpythongo
0
3
Eci EvalA

This evaluation protocol assesses the capability of NLP models to identify causal relationships between event pairs in text. It probes both sentence-level and document-level reasoning, measuring how well models can distinguish true causal links from mere correlations or coincidental co-occurrences. Use when the user wants to benchmark on CTB, ESL, MAVEN-ERE, MECI, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3