Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,665–5,688 of 23,785 skills
Evaluates how well automated metrics and vision-language models can identify granular errors (attribute/relation misattachments) in detailed image descriptions and correctly rank paired descriptions against human judgments. Use when the user has predictions and gold and needs to compute macro F1, Spearman rank ρ.
Evaluates a model's ability to estimate 6-DOF camera pose (translation and rotation) from a single monocular image across indoor and outdoor environments. It probes the network's robustness to challenging conditions like motion blur, low light, and dynamic objects, as well as its generalization to unseen scenes and varying training baselines. Use when the user wants to benchmark on 7 Scenes, Cambridge Landmarks, or asks about evaluating this task. Reports localization error.
Evaluates the robustness of human and animal pose estimation models when subjected to real-world image corruptions such as blur, noise, compression, lighting changes, and occlusion masks. It measures how much model accuracy degrades relative to clean-image performance across varying corruption severities. Use when the user wants to benchmark on COCO-C, OCHuman-C, AP10K-C, or asks about evaluating this task. Reports mRR.
Evaluates the ability of self-supervised visual representations to capture geometric pose information and semantic content. It probes absolute and relative pose estimation accuracy, as well as semantic classification performance, across in-domain, out-of-domain, and real-world settings. Use when the user wants to benchmark on Carvana, Synthetic dataset [8], or asks about evaluating this task. Reports relative pose estimation accuracy.
Evaluates neural models on a Portuguese-language benchmark derived from English GLUE and SuperGLUE tasks, probing capabilities in grammatical acceptability, sentiment, paraphrase detection, semantic similarity, natural language inference, reading comprehension, and causal reasoning. Use when the user wants to benchmark on CoLA, SST-2, MRPC, QQP, STS-B, WiC, MNLI, QNLI, RTE, WNLI, WSC, CB, AXb, AXg, BoolQ, MultiRC, ReCoRD, COPA, or asks about evaluating this task. Reports single-number perform...
Evaluates multimodal models on portrait composition understanding and generation. It probes the ability to predict aesthetic scores, reason about fine-grained composition attributes, answer image-grounded questions, and generate portraits that adhere to explicit spatial and compositional constraints. Use when the user wants to benchmark on PortraitCraft, or asks about evaluating this task. Reports SRCC.
Evaluates large language models' ability to perform quantitative reasoning and structured decision-making in financial portfolio optimization. It probes whether models can correctly apply convex optimization principles under varying constraints and multi-criteria objectives. Use when the user wants to benchmark on PortBench, or asks about evaluating this task. Reports accuracy.
This evaluation probes the robustness of vision-language computer agents against adversarial visual distractions (pop-ups) injected into GUI environments. It measures how often agents are tricked into interacting with malicious overlays and how these distractions degrade their ability to complete legitimate user tasks. Use when the user wants to benchmark on OSWorld, VisualWebArena, or asks about evaluating this task. Reports Attack Success Rate (ASR).
Tests object perception and hallucination on images without captions, evaluating whether LVLMs can ground object detection purely from visual input without textual priors. Use when the user wants to benchmark on POPE-NoCaps, or asks about evaluating this task. Reports Acc.
Evaluates object perception and hallucination in LVLMs by prompting models to identify whether specific objects are present in an image. It measures how often models correctly affirm or deny object existence without generating false positives. Use when the user wants to benchmark on POPE, or asks about evaluating this task. Reports Acc.
This benchmark evaluates the effectiveness of feature augmentation modules for detecting Ponzi scheme accounts on the Ethereum blockchain. It probes a model's ability to classify account nodes as legitimate or malicious based on transaction graph structures and temporal behavior patterns. Use when the user wants to benchmark on Ethereum Ponzi dataset, or asks about evaluating this task. Reports micro-F1.
Evaluates a unified foundation model's ability to perform joint polyp detection, segmentation, classification, and unsupervised tracking on colonoscopy video frames. It tests generalization to unseen clinical datasets and consistency of object association across frames without task-specific fine-tuning. Use when the user wants to benchmark on Kvasir-SEG, CVC-ClinicDB, CVC-ColonDB, ETIS, CVC-300, KUMC, REAL-Colon, or asks about evaluating this task. Reports Dice.
Evaluates the accuracy and continuity of polyploid haplotype assembly methods by measuring switch errors and read-haplotype conflicts, while also quantifying phasing uncertainty across varying ploidies, coverages, and genomic structures. Use when the user wants to benchmark on Synthetic Polyploid Genomes (S. tuberosum), Experimental Octoploid Strawberry (F. x ananassa), or asks about evaluating this task. Reports Generalized Vector Error Rate (VER).
Evaluates a model's ability to predict polypharmacy side effects (drug-drug interactions) in a multimodal biomedical graph. It probes the model's capacity to learn continuous latent representations for drugs and proteins and generalize to unseen drug pairs across 964 specific side effect types. Use when the user wants to benchmark on Polypharmacy Side Effects Dataset, or asks about evaluating this task. Reports cross-entropy loss.
Evaluates multi-modal mathematical and cognitive reasoning capabilities on visual puzzles. It probes spatial interpretation, relational understanding, pattern recognition, and long-horizon logical reasoning using diagram-based multiple-choice questions. Use when the user wants to benchmark on POLYMATH, or asks about evaluating this task. Reports accuracy.
Evaluates the toxicity of LLM-generated continuations across 17 languages using naturally occurring prompts scraped from the web. It probes how model size, language resource availability, and instruction/preference tuning affect the generation of harmful content. Use when the user wants to benchmark on PolygloToxicityPrompts (PTP), or asks about evaluating this task. Reports AT.
Evaluates large language models' ability to detect hallucinations by verifying factual claims across 11 languages. It probes cross-linguistic consistency, topic-aware fact-checking, and resistance to web-resource bias in multilingual settings. Use when the user wants to benchmark on Poly-FEVER, or asks about evaluating this task. Reports accuracy.
Evaluates open-domain question answering in Polish by measuring both passage retrieval accuracy and answer generation quality. It probes a model's ability to retrieve relevant evidence from a large corpus and accurately extract or generate answers from those passages. Use when the user wants to benchmark on PolQA, or asks about evaluating this task. Reports fuzzy_match.
Evaluates the alignment and quality of generated image captions relative to reference captions and source images. It measures how well a learned metric correlates with human judgments, specifically probing hallucination robustness and open-vocabulary caption evaluation. Use when the user has predictions and gold and needs to compute Polos.
Evaluates the ability of LLMs and API-based classifiers to accurately annotate toxicity and incivility in political protest content against a human gold standard. It probes zero-shot classification performance, threshold sensitivity, and output reproducibility across different model sizes and temperatures. Use when the user wants to benchmark on Political protest content dataset, or asks about evaluating this task. Reports F1-score.
Evaluates Polish language understanding, summarization, and question answering capabilities of text-to-text models. It probes how well encoder-decoder and decoder-only architectures generalize from multilingual pre-training to monolingual Polish tasks using exact-match generation and ROUGE-based metrics. Use when the user wants to benchmark on KLEJ benchmark, Allegro Articles, Polish Summaries Corpus, or asks about evaluating this task. Reports exact-match accuracy.
Evaluates large language models on Polish medical licensing and specialization exams to assess cross-lingual medical knowledge transfer, domain-specific understanding, and specialty-level accuracy compared to human medical graduates. Use when the user wants to benchmark on Polish Medical Exams (LEK/LDEK/PES), or asks about evaluating this task. Reports score.
Evaluates the transcription accuracy of various automatic speech recognition (ASR) models on Polish-language audio, contrasting read-speech benchmarks with spontaneous, noisy medical consultations to probe domain generalization. Use when the user wants to benchmark on Mozilla Common Voice (MCV) Polish, Multilingual LibriSpeech (MLS) Polish, Medical interview corpus, or asks about evaluating this task. Reports Word Error Rate (WER).
Evaluates the model's ability to retrieve relevant coordination policies that guide task planning based on a high-level progress summary of the current state. Use when the user wants to benchmark on Policy Selection Test Suite, or asks about evaluating this task. Reports F1 Score.