Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,657–7,680 of 20,850 skills
Evaluates the ability of NLP models to detect and classify hate speech in social media comments. It probes both binary hate/non-hate classification and multi-label categorization across specific demographic/identity-based hate categories. The benchmark emphasizes handling overlapping labels and class imbalance typical of real-world user-generated text. Use when the user wants to benchmark on ETHOS, or asks about evaluating this task. Reports F1-score (macro).
Machine translation performance across multiple low-resource Ethiopian languages paired with English, evaluating both English-to-Ethiopian and Ethiopian-to-English directions. It probes how model initialization (training from scratch vs. fine-tuning a multilingual model) and available corpus size impact translation quality. Use when the user wants to benchmark on EthioMT, or asks about evaluating this task. Reports spBLEU.
Evaluates LLMs' ability to classify multiple emotions in low-resource Ethiopian languages (Amharic, Afan Oromo, Somali, Tigrinya) and English. It probes cross-lingual transfer capabilities, the effectiveness of zero-shot and few-shot prompting strategies, and the impact of fine-tuning on multi-label emotion understanding tasks. Use when the user wants to benchmark on EthioEmo, or asks about evaluating this task. Reports Weighted-averaged F1-score.
Evaluates a model's ability to detect fraudulent Ethereum accounts by analyzing transaction interaction graphs and behavioral patterns. It specifically probes the model's capacity to handle imbalanced label distributions and distinguish between Ponzi schemes and phishing scams using self-supervised feature learning. Use when the user wants to benchmark on Ponzi Scheme Dataset, Phish Scam Dataset, or asks about evaluating this task. Reports Binary-F1.
Evaluates the psychological and behavioral impact of an AI-generated emotional self-voice intervention compared to text-only and control conditions on goal-related resilience, confidence, motivation, and emotional states. Use when the user wants to benchmark on Custom Human-Subject Intervention Dataset, or asks about evaluating this task. Reports Self-report questionnaire scores.
Evaluates large language models on native Estonian language capabilities across seven tasks covering factual recall, grammar, morphology, vocabulary, summarization, and structured information extraction. The benchmark emphasizes cultural and linguistic authenticity by using human-curated or native-source data without machine translation, assessing both general and domain-specific competencies. Use when the user wants to benchmark on Exams, Trivia, Declension, Words, Grammar, News, Speaker Nam...
Evaluates the quality and linguistic characteristics of argumentative essays generated by different AI models compared to human-written texts. It probes logical structure, vocabulary richness, syntactic complexity, and stylistic markers through expert human annotation. Use when the user wants to benchmark on Student Essay Dataset (90 topics), or asks about evaluating this task. Reports Mean Rating Score.
Evaluates the robustness of AI-generated text detectors against back-translation manipulations. It probes whether detectors can maintain high true positive rates when AI-generated text is translated to intermediate languages and back-translated to English, preserving semantics while evading detection. Use when the user wants to benchmark on ESPERANTO, or asks about evaluating this task. Reports True Positive Rate (TPR).
Evaluates multi-task learning architectures for post-click conversion rate (CVR) prediction in online advertising, specifically testing how parameter sharing and entire-space training mitigate data sparsity and bias. Use when the user wants to benchmark on Unspecified industry advertising dataset, or asks about evaluating this task. Reports Performance.
Evaluates a model's ability to estimate post-click conversion rate (CVR) and post-click-and-conversion rate (CTCVR) in recommendation systems. It specifically probes how well the model handles sample selection bias and data sparsity by comparing performance on clicked-only impressions versus the entire impression space. Use when the user wants to benchmark on Public Dataset, or asks about evaluating this task. Reports AUC.
Evaluates e-commerce language understanding through masked token recovery on product texts and graded semantic similarity between search queries and products. Also assesses general natural language understanding capabilities via the GLUE benchmark. Use when the user wants to benchmark on Amazon ESCI, GLUE, or asks about evaluating this task. Reports top-k accuracy, Spearman correlation.
Evaluates zero-shot language-queried audio source separation on environmental sound classes. The benchmark tests the model's ability to isolate a target sound from a mixed audio mixture using text labels. Use when the user wants to benchmark on ESC-50, or asks about evaluating this task. Reports SDRi.
This benchmark evaluates the clinical readiness of Large Vision Language Models (LVLMs) for emergency room monitoring tasks. It probes their ability to generate accurate, clinically cautious, and semantically entailed long-form answers from medical images, while identifying specific failure modes like hallucinations and overconfidence. Use when the user wants to benchmark on ERVQA, or asks about evaluating this task. Reports Entailment Score.
This benchmark evaluates the environmental resilience of discrete speech codecs by measuring how reconstruction quality and downstream task performance degrade under varying signal-to-noise ratios, loudness levels, and real-world acoustic conditions. It probes both signal fidelity and semantic/intelligibility consistency after codec compression and subsequent speech enhancement or recognition. Use when the user wants to benchmark on Environment-Resilient Speech Codec Benchmark (ERSB), or asks...
Evaluates a model's ability to detect classification errors and recover hierarchical multi-label constraints without prior knowledge. It probes the system's capacity to generate interpretable logical rules from failure patterns and improve downstream model consistency. Use when the user wants to benchmark on Military Vehicles, ImageNet50, OpenImage36, or asks about evaluating this task. Reports F1-score.
Evaluates a trillion-parameter multimodal foundation model across text, vision, audio, and generation tasks to measure factual knowledge, reasoning, coding, instruction following, and agent capabilities. Use when the user wants to benchmark on PreciseWikiQA, MMLU-Pro, MATH, LiveCodeBench, MMMU-Pro, MathVista, GenEval, VBench, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of spatiotemporal deep learning models to classify ocular ultrasound videos into binary diagnostic categories: detecting retinal detachment (RD) versus non-RD, and classifying macular status (intact vs. detached). It probes the model's capacity to learn subtle spatiotemporal patterns in medical ultrasound while handling real-world class imbalance. Use when the user wants to benchmark on ERDES, or asks about evaluating this task. Reports F1-Score.
Evaluates a model's ability to recognize and classify emotional states from conversational dialogue context. It probes multi-turn emotional understanding, speaker-aware reasoning, and generalization across diverse domain settings and speaker demographics. Use when the user wants to benchmark on IEMOCAP, MELD, EmoryNLP, or asks about evaluating this task. Reports Accuracy.
Evaluates NLP models' ability to generate faithful, task-appropriate rationales for predictions, measuring both alignment with human annotations and causal faithfulness via token perturbation. Use when the user wants to benchmark on Movies, FEVER, CoS-E, eSNLI, or asks about evaluating this task. Reports AUPRC.
Evaluates machine unlearning algorithms in recommender systems across collaborative filtering, session-based, and next-basket recommendation tasks. It measures how well models retain recommendation quality after unlearning sensitive or malicious user interactions, while maintaining computational efficiency and effectiveness compared to full retraining. Use when the user wants to benchmark on ERASE Benchmark (9 datasets), or asks about evaluating this task. Reports utility.
Spatio-temporal forecasting of atmospheric temperature using deep learning models. It probes the ability to capture long-range spatial-temporal dependencies and predict future weather states from historical multi-feature sequences. Use when the user wants to benchmark on ERA5 Turkey, WeatherBench, or asks about evaluating this task. Reports RMSE.
Evaluates LLMs on longitudinal clinical reasoning across five emergency room workflow stages, including acuity assessment, case summarization, treatment planning, final diagnosis, and patient disposition. It probes the models' ability to integrate sparse clinical notes, perform rule-out differential diagnosis, and align outputs with real-world clinical decision-making and safety constraints. Use when the user wants to benchmark on ER-Reason, or asks about evaluating this task. Reports Accurac...
Evaluates large language models' ability to understand and rate emotional intensity in conflict-driven dialogue scenarios. It probes emotional intelligence through automated scoring of model-generated ratings, avoiding subjective human interpretation. Use when the user wants to benchmark on EQ-Bench, or asks about evaluating this task. Reports EQ-Bench Score.
Evaluates the accuracy of AI weather forecasting models against established numerical models and ground-truth observations. It probes the model's ability to predict atmospheric variables (e.g., wind speed, solar radiation) at hourly resolution over 20-day lead times. Use when the user wants to benchmark on ERA5, IFS HRES IC, Weather Stations, or asks about evaluating this task. Reports Skill Score (SS).