Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,849
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,681–7,704 of 20,849 skills

Ept Benchmark EvalA

Assesses the trustworthiness of large language models within a Persian-Islamic cultural context across six dimensions: Ethics, Fairness, Privacy, Robustness, Safety, and Truthfulness. It measures how well model responses align with culturally specific ethical principles and linguistic nuances. Use when the user wants to benchmark on EPT Benchmark, or asks about evaluating this task. Reports compliance metric.

researchpythonrust
0
3
Epsos Nids EvalA

Evaluates the ability of machine learning and deep learning classifiers, particularly Decision Trees optimized with Enhanced Particle Swarm Optimization, to accurately detect and classify multiple types of network intrusions in high-dimensional traffic data. Use when the user wants to benchmark on CSE-CIC-IDS-2018, LITNET-2020, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
EpsilonA

Evaluates the correlation between a zero-cost NAS metric (epsilon) and actual training accuracy across different neural architecture search spaces, testing the metric's ability to rank architectures without training. It probes whether output dispersion from constant weight initializations can serve as a reliable, hyperparameter-free proxy for architecture selection. Use when the user has predictions and gold and needs to compute Spearman ρ (global).

researchpythongo
0
3
Epidemiological Benchmark EvalA

Evaluates spatio-temporal graph models on forecasting, stability, and denoising tasks using synthetic epidemiological data generated from PDEs. Probes the model's ability to predict future states, resist noise/dropout, and recover clean signals from corrupted graph time-series. Use when the user wants to benchmark on Epidemiological Synthetic Dataset, or asks about evaluating this task. Reports RMSE.

researchpythonnode
0
3
Epic Kitchens 100 Mqa EvalA

Evaluates multi-modal large language models' ability to recognize and distinguish between similar human actions in egocentric videos through multiple-choice question answering. It specifically probes fine-grained action discrimination using hard, semantically and visually similar distractors generated by action recognition models. Use when the user wants to benchmark on EPIC-KITCHENS-100-MQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Epd EvalA

Evaluates how well image quality assessment models correlate with actual robotic task performance under various image distortions. It probes whether traditional human-centric visual quality metrics align with the perception needs of embodied robots performing push and pick tasks. Use when the user wants to benchmark on EPD, or asks about evaluating this task. Reports PLCC.

researchpythongo
0
3
Eo Bench EvalA

Evaluates a model's ability to reason about embodied interactions, including spatial understanding, physical commonsense, task planning, and state estimation from robot vision and text inputs. Use when the user wants to benchmark on EO-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Enviroexam EvalA

This benchmark evaluates large language models' domain-specific knowledge in environmental science using multiple-choice questions derived from university curricula. It measures both raw accuracy and performance consistency across different course topics, revealing how well models retain and apply specialized scientific concepts. Use when the user wants to benchmark on EnviroExam, or asks about evaluating this task. Reports composite_index.

researchpythongo
0
3
Entropy Minimization Reasoning EvalA

Evaluates the reasoning capabilities of LLMs on mathematical and coding benchmarks by measuring accuracy under various inference-time scaling and unsupervised entropy minimization techniques. It probes whether reducing output uncertainty improves correctness without labeled data or parameter updates. Use when the user wants to benchmark on AMC, AIME, Minerva, LeetCode Live Contest, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Entropy Adaptive Merging EvalA

Evaluates the robustness of an entropy-adaptive model merging framework under heterogeneous domain shifts. It probes whether a merged model can adapt to unseen target domains using only forward passes and entropy-based coefficients, without backpropagation or labeled target data. Use when the user wants to benchmark on MiDog Atypical, Organs, Histopantum, ISIC Skin, Messidor, PACS, VLCS, Office-Home, TerraIncognita, or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Entity6k EvalA

Evaluates open-domain entity recognition capabilities across four visual grounding and understanding tasks: object detection, zero-shot image classification, image captioning, and dense captioning. It measures how well models can localize, classify, and describe specific real-world entities in images without fine-tuning. Use when the user wants to benchmark on Entity6K, or asks about evaluating this task. Reports AP.

researchpythongo
0
3
Entity Linking EvalA

Evaluates end-to-end entity linking systems on their ability to detect entity mentions and correctly disambiguate them to knowledge base entities. It specifically probes for systemic benchmark biases, such as overreliance on named entities, ambiguous disambiguation choices, and underrepresented entity types, by introducing fairer evaluation protocols. Use when the user wants to benchmark on Existing and new EL benchmarks, or asks about evaluating this task. Reports Micro F1.

researchpythongo
0
3
Entity Hallucination EvalA

Evaluates large language models' ability to generate factually correct answers to complex and factual questions while mitigating entity-level hallucinations. It also measures the effectiveness of a real-time hallucination detection mechanism in identifying fabricated or low-confidence entities during generation. Use when the user wants to benchmark on WikiBio GPT-3 dataset, 2WikiMultihopQA, StrategyQA, NQ, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Entity Canonicalization EvalA

Evaluates a model's ability to cluster entity mentions (noun phrases) into canonical forms within open knowledge graphs. It probes unsupervised representation learning and clustering capabilities without relying on manually annotated ground truth for training. Use when the user wants to benchmark on Base, Ambiguous, ReVerb45K, CanonicNell, or asks about evaluating this task. Reports Macro F1.

researchpythongo
0
3
Entities Of The Union EvalA

Evaluates the ability of models to disambiguate named entity mentions in historical and modern newswire texts, and to cluster coreferent mentions across documents. It specifically probes handling of out-of-knowledgebase individuals common in historical contexts. Use when the user wants to benchmark on Entities of the Union, MSNBC, ACE2004, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Enterprise Sql Kg Qa EvalA

Evaluates large language models' ability to generate correct SQL or SPARQL queries for natural language questions over an enterprise insurance database. It measures how well zero-shot prompting with raw schema versus knowledge graph augmentation improves factual grounding and query execution accuracy. Use when the user wants to benchmark on Enterprise SQL & KG QA Benchmark, or asks about evaluating this task. Reports execution accuracy.

researchpythonsql
0
3
Enterprise Benchmarks EvalA

Evaluates LLMs on domain-specific enterprise tasks across finance, legal, climate, and cybersecurity. It probes capabilities like numerical reasoning, named entity recognition, document relevance ranking, and long-document summarization using real-world industry data. Use when the user wants to benchmark on Earnings Call Transcripts, News Headline, Credit Risk Assessment (NER), KPI-Edgar, FiNER-139, Opinion-based QA (FiQA), Sentiment Analysis (FiQA SA), Insurance QA, ConvFinQA, Financial Text...

researchpythongo
0
3
Enginead EvalA

Probes the ability of one-class anomaly detection algorithms to identify incipient engine faults in real-world, multivariate vehicle sensor telemetry. It specifically evaluates cross-vehicle generalization and robustness to distributional shifts in normal operating conditions across a commercial fleet. Use when the user wants to benchmark on EngineAD, or asks about evaluating this task. Reports F1-score (anomaly class).

researchpythongo
0
3
Energy First Arch EvalA

Evaluates the classification accuracy and training energy efficiency of biologically-inspired and physics-guided neural architectures against conventional baselines across diverse data modalities. It probes whether action-principle regularization yields modality-specific performance gains and reduced internal activation energy without accuracy loss. Use when the user wants to benchmark on Fashion-MNIST, CIFAR-10, DVS Gesture, SHD, SSC, WESAD, DREAMER, SEED-IV, 20newsgroups, or asks about eval...

researchpythongo
0
3
Energy EfficiencyA

Evaluates the trade-off between energy efficiency and physical-layer security in an IRS-assisted MISO network with cooperative jamming. Probes how varying transmit power constraints and secrecy rate thresholds impact system performance compared to baseline beamforming strategies. Use when the user has predictions and gold and needs to compute Energy Efficiency.

researchpythongo
0
3
Energaizer Gpu Power EvalA

Evaluates the accuracy of a lightweight analytical framework in predicting GPU latency and dynamic power consumption for AI workloads across different hardware architectures, operating frequencies, and algorithm configurations. Use when the user wants to benchmark on EnergAIzer Kernel Database & AI Workloads, or asks about evaluating this task. Reports MAPE.

researchpythongo
0
3
Enem Vision EvalA

Evaluates multimodal and text-only language models on Brazilian university admission exams (ENEM), specifically probing their ability to comprehend visual information, interpret tables/figures, and perform mathematical reasoning in a multiple-choice format. Use when the user wants to benchmark on ENEM 2022/2023, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Endo Depth Robustness EvalA

This benchmark evaluates the robustness of monocular depth estimation models when processing endoscopic images degraded by realistic surgical artifacts. It probes how well models maintain depth prediction accuracy and consistency under varying severities of illumination changes, optical blurs, visual obstructions, sensor noise, and compression artifacts. Use when the user wants to benchmark on Endoscopic Depth Estimation Dataset (Synthetically Corrupted), or asks about evaluating this task. R...

researchpythongit
0
3
Encqa EvalA

Evaluates vision-language models on their ability to interpret visual encodings (position, length, area, color, shape) and perform chart-specific analytic tasks (e.g., value retrieval, anomaly detection, correlation estimation). It probes fine-grained visual perception, reasoning under different encoding constraints, and whether model capabilities scale with size or prompting strategies. Use when the user wants to benchmark on EncQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3