Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,827
skills in category
868
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,841–6,864 of 20,827 skills

Iirc EvalA

Evaluates lifelong learning algorithms on incremental label refinement, requiring models to predict both coarse (superclass) and fine-grained (subclass) labels over time without forgetting prior knowledge, while operating under incomplete information constraints. Use when the user wants to benchmark on IIRC-CIFAR, IIRC-ImageNet, or asks about evaluating this task. Reports pw-JS.

researchpythongo
0
3
Igenbench EvalA

Probes the reliability of text-to-infographic generation models by decomposing visual fidelity into atomic yes/no checks. It evaluates whether generated images accurately encode data, follow structural constraints, and maintain consistency across multiple verification questions. Use when the user wants to benchmark on IGenBench, or asks about evaluating this task. Reports Q-ACC.

researchpython
0
3
Igbo English Mt EvalA

This benchmark evaluates bidirectional machine translation quality between Igbo and English. It probes a model's ability to accurately translate news and contemporary media content across two directions (Igbo-to-English and English-to-Igbo) using human-validated parallel sentences. Use when the user wants to benchmark on Igbo-English MT Benchmark, or asks about evaluating this task. Reports BLEU.

researchpythongit
0
3
Ifir EvalA

This benchmark evaluates an information retrieval system's ability to follow complex, domain-specific instructions when retrieving relevant passages. It probes whether models can interpret nuanced constraints (e.g., patient demographics, legal case details, financial goals) rather than just matching keyword semantics. Use when the user wants to benchmark on IfIR, or asks about evaluating this task. Reports InstFol@20.

researchpythongo
0
3
Ifeval EvalA

Evaluates large language models' ability to follow explicit, verifiable instructions embedded in prompts, such as length constraints, keyword inclusion, formatting rules, and language requirements. It measures both strict and loose compliance across individual instructions and entire prompts to assess structural and syntactic robustness. Use when the user wants to benchmark on IFEval, or asks about evaluating this task. Reports Inst-level strict-accuracy.

researchpythongo
0
3
Ifc Bench V2 EvalA

Tests an LLM's ability to extract, compute, and reason over heterogeneous Building Information Modeling (BIM) data (IFC files) using adaptive code execution or static baselines. It probes robustness to data heterogeneity, documentation retrieval, and tool augmentation. Use when the user wants to benchmark on ifc-bench v2, or asks about evaluating this task. Reports aggregate accuracy.

researchpythongo
0
3
Iemocap EvalA

Evaluates how categorical and continuous label ambiguity impacts the performance of unimodal emotion recognition models (text, audio, facial) on the IEMOCAP dataset. It tests whether filtering data by annotator agreement or VAD score dispersion yields cleaner evaluation signals. The protocol highlights the disconnect between rigid single-label benchmarks and the inherent ambiguity of affective data. Use when the user wants to benchmark on IEMOCAP, or asks about evaluating this task. Reports w...

researchpythongo
0
3
Ielm EvalA

Evaluates the zero-shot open information extraction (OIE) capability of pre-trained language models by measuring their ability to extract subject-predicate-object triples from text without task-specific training or fine-tuning. It probes whether LMs inherently store rich, open-world relational knowledge that can be accessed via attention mechanisms. Use when the user wants to benchmark on CaRB, Re-OIE2016, TAC KBP-OIE, Wikidata-OIE, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Ieee Cis Fraud EvalA

Evaluates binary financial fraud detection performance across diverse model architectures (LSTM, Transformer, XGBoost, GNN, and ensembles) on highly imbalanced transaction data. Probes threshold-independent discrimination (AUC-ROC, PR-AUC) and threshold-dependent detection accuracy (F1, Precision, Recall, MCC) under stratified cross-validation and temporal holdout conditions. Use when the user wants to benchmark on IEEE-CIS Financial Fraud Detection Dataset, or asks about evaluating this task...

researchpythonperformance
0
3
Idsr EvalA

Evaluates the accuracy and diversity of end-to-end sequential recommendation models by testing their ability to predict the next item in a user's behavior sequence while maintaining item diversity in the recommendation list. Use when the user wants to benchmark on ML100K, ML1M, or asks about evaluating this task. Reports Recall.

researchpythontesting
0
3
Ids Smart Grid EvalA

This benchmark evaluates machine learning-based anomaly detection systems for smart grid cybersecurity, focusing on both detection performance and model explainability. It probes how well intrusion detection methods generalize across diverse operational datasets while providing interpretable feature importance and robustness to data noise. Use when the user wants to benchmark on Power System dataset, CIDDS-002 dataset, or asks about evaluating this task. Reports explanation sensitivity (Expl....

researchpythonrust
0
3
Ids Moo Automl EvalA

Evaluates intrusion detection systems for resource-constrained IoT and cloud environments by measuring classification accuracy, computational efficiency, and model confidence. It probes the ability of AutoML pipelines to balance detection performance against training time, inference latency, and memory footprint. Use when the user wants to benchmark on CICIDS2017, IoTID20, or asks about evaluating this task. Reports F1-score.

researchpythonsecurity
0
3
Ids Detection EvalA

Evaluates machine learning classifiers for network intrusion detection on imbalanced, high-dimensional traffic data. Probes the model's ability to distinguish benign from malicious traffic across binary and multilabel settings using standard classification metrics. Use when the user wants to benchmark on UNSW-NB15, CIC-IDS2017, CIC-IDS2018, or asks about evaluating this task. Reports Accuracy.

researchpythonperformance
0
3
Ids Automl EvalA

Evaluates an AutoML-based intrusion detection system's ability to classify network traffic as benign or malicious across multiple attack types. It probes the framework's robustness to class imbalance and its efficiency in real-time network environments. Use when the user wants to benchmark on CICIDS2017, 5G-NIDD, or asks about evaluating this task. Reports F1-score.

researchpythonsecurity
0
3
Idnet Dataset EvalA

Evaluates the quality and utility of a large-scale synthetic identity document dataset for fraud detection. It measures metadata diversity, visual fidelity to real documents, stealthiness of forged modifications, and downstream model accuracy. Use when the user wants to benchmark on IDNet, or asks about evaluating this task. Reports SSIM.

researchpythonperformance
0
3
Idiomaticity Detection EvalA

Evaluates large language models' ability to disambiguate whether a given phrase is used idiomatically or literally within a specific context. It probes zero-shot, few-shot, and cross-lingual prompting capabilities, measuring how well models generalize to idiomatic expressions without task-specific fine-tuning. Use when the user wants to benchmark on SemEval 2022 Task 2a, FLUTE, MAGPIE, or asks about evaluating this task. Reports macro F1.

researchpythongo
0
3
Idh Mutation Prediction EvalA

Evaluates a model's ability to predict IDH mutation status (mutant vs. wild-type) in glioma patients using multi-modal MRI-derived structural brain networks. Use when the user wants to benchmark on TCIA + In-house Cohort, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Identity Fraud Detection EvalA

This evaluation probes a dialogue system's ability to dynamically generate derived questions and manage multi-turn interactions to accurately classify loan applicants as fraudulent or legitimate based on their knowledge of personal information triplets. Use when the user wants to benchmark on Applicant Personal Information Dataset, or asks about evaluating this task. Reports recognition accuracy.

researchpythongo
0
3
Ideation Space EvalA

Evaluates a framework's ability to decompose scientific papers into orthogonal conceptual dimensions (problem, method, findings) and model transitions between them. It probes fine-grained conceptual similarity retrieval and assesses whether the model's novelty predictions align with expert human judgments. Use when the user wants to benchmark on ICLR 2025 Submissions, AI-Researcher, or asks about evaluating this task. Reports Recall@K.

researchpythongo
0
3
Ideabench EvalA

Evaluates the professional design capabilities of generative models across text-to-image, image-to-image, and multi-image generation tasks. It probes aesthetic quality, contextual relevance, multimodal alignment, and adherence to complex, real-world design requirements that go beyond basic generation. Use when the user wants to benchmark on IDEA-Bench, or asks about evaluating this task. Reports Avg. Score.

researchpythongo
0
3
Idd Aw EvalA

Evaluates the robustness and safety of semantic segmentation models for autonomous driving in unstructured traffic and adverse weather. It specifically probes whether models can correctly identify critical road elements and traffic participants when visual quality degrades due to rain, fog, snow, or low light. Use when the user wants to benchmark on IDD-AW, or asks about evaluating this task. Reports Safe mIoU (SmIoU).

researchpythongo
0
3
Icu Readmission Prediction EvalA

Predicts the risk of a patient being readmitted to the ICU within 30 days of discharge using longitudinal electronic medical record (EMR) data. The task evaluates how well different deep learning architectures can model time-varying clinical events (diagnoses, procedures, medications, vital signs) alongside static demographic covariates to capture complex patient trajectories and risk factors. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports AUROC.

researchpythongit
0
3
Icrl Molecular EvalA

Evaluates whether text-based LLMs can effectively leverage high-dimensional non-text modality representations (e.g., molecular embeddings from foundation models) via training-free in-context learning, comparing various representation injection and projection strategies. Use when the user wants to benchmark on ESOL, Caco_wang, AqSolDB, LD50_Zhu, AstraZeneca, or asks about evaluating this task. Reports RMSE.

researchpythongit
0
3
Iclr2025 Review Feedback EvalA

Evaluates the causal impact of LLM-generated feedback on peer review quality by measuring reviewer engagement, revision rates, feedback incorporation, and downstream rebuttal dynamics in a large-scale randomized controlled trial. Use when the user wants to benchmark on ICLR 2025 Reviews, or asks about evaluating this task. Reports update_rate.

researchpythongit
0
3