Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,617–5,640 of 23,734 skills
Evaluates object perception and hallucination in LVLMs by prompting models to identify whether specific objects are present in an image. It measures how often models correctly affirm or deny object existence without generating false positives. Use when the user wants to benchmark on POPE, or asks about evaluating this task. Reports Acc.
This benchmark evaluates the effectiveness of feature augmentation modules for detecting Ponzi scheme accounts on the Ethereum blockchain. It probes a model's ability to classify account nodes as legitimate or malicious based on transaction graph structures and temporal behavior patterns. Use when the user wants to benchmark on Ethereum Ponzi dataset, or asks about evaluating this task. Reports micro-F1.
Evaluates a unified foundation model's ability to perform joint polyp detection, segmentation, classification, and unsupervised tracking on colonoscopy video frames. It tests generalization to unseen clinical datasets and consistency of object association across frames without task-specific fine-tuning. Use when the user wants to benchmark on Kvasir-SEG, CVC-ClinicDB, CVC-ColonDB, ETIS, CVC-300, KUMC, REAL-Colon, or asks about evaluating this task. Reports Dice.
Evaluates the accuracy and continuity of polyploid haplotype assembly methods by measuring switch errors and read-haplotype conflicts, while also quantifying phasing uncertainty across varying ploidies, coverages, and genomic structures. Use when the user wants to benchmark on Synthetic Polyploid Genomes (S. tuberosum), Experimental Octoploid Strawberry (F. x ananassa), or asks about evaluating this task. Reports Generalized Vector Error Rate (VER).
Evaluates a model's ability to predict polypharmacy side effects (drug-drug interactions) in a multimodal biomedical graph. It probes the model's capacity to learn continuous latent representations for drugs and proteins and generalize to unseen drug pairs across 964 specific side effect types. Use when the user wants to benchmark on Polypharmacy Side Effects Dataset, or asks about evaluating this task. Reports cross-entropy loss.
Evaluates multi-modal mathematical and cognitive reasoning capabilities on visual puzzles. It probes spatial interpretation, relational understanding, pattern recognition, and long-horizon logical reasoning using diagram-based multiple-choice questions. Use when the user wants to benchmark on POLYMATH, or asks about evaluating this task. Reports accuracy.
Evaluates the toxicity of LLM-generated continuations across 17 languages using naturally occurring prompts scraped from the web. It probes how model size, language resource availability, and instruction/preference tuning affect the generation of harmful content. Use when the user wants to benchmark on PolygloToxicityPrompts (PTP), or asks about evaluating this task. Reports AT.
Evaluates large language models' ability to detect hallucinations by verifying factual claims across 11 languages. It probes cross-linguistic consistency, topic-aware fact-checking, and resistance to web-resource bias in multilingual settings. Use when the user wants to benchmark on Poly-FEVER, or asks about evaluating this task. Reports accuracy.
Evaluates open-domain question answering in Polish by measuring both passage retrieval accuracy and answer generation quality. It probes a model's ability to retrieve relevant evidence from a large corpus and accurately extract or generate answers from those passages. Use when the user wants to benchmark on PolQA, or asks about evaluating this task. Reports fuzzy_match.
Evaluates the alignment and quality of generated image captions relative to reference captions and source images. It measures how well a learned metric correlates with human judgments, specifically probing hallucination robustness and open-vocabulary caption evaluation. Use when the user has predictions and gold and needs to compute Polos.
Evaluates the ability of LLMs and API-based classifiers to accurately annotate toxicity and incivility in political protest content against a human gold standard. It probes zero-shot classification performance, threshold sensitivity, and output reproducibility across different model sizes and temperatures. Use when the user wants to benchmark on Political protest content dataset, or asks about evaluating this task. Reports F1-score.
Evaluates Polish language understanding, summarization, and question answering capabilities of text-to-text models. It probes how well encoder-decoder and decoder-only architectures generalize from multilingual pre-training to monolingual Polish tasks using exact-match generation and ROUGE-based metrics. Use when the user wants to benchmark on KLEJ benchmark, Allegro Articles, Polish Summaries Corpus, or asks about evaluating this task. Reports exact-match accuracy.
Evaluates large language models on Polish medical licensing and specialization exams to assess cross-lingual medical knowledge transfer, domain-specific understanding, and specialty-level accuracy compared to human medical graduates. Use when the user wants to benchmark on Polish Medical Exams (LEK/LDEK/PES), or asks about evaluating this task. Reports score.
Evaluates the transcription accuracy of various automatic speech recognition (ASR) models on Polish-language audio, contrasting read-speech benchmarks with spontaneous, noisy medical consultations to probe domain generalization. Use when the user wants to benchmark on Mozilla Common Voice (MCV) Polish, Multilingual LibriSpeech (MLS) Polish, Medical interview corpus, or asks about evaluating this task. Reports Word Error Rate (WER).
Evaluates the model's ability to retrieve relevant coordination policies that guide task planning based on a high-level progress summary of the current state. Use when the user wants to benchmark on Policy Selection Test Suite, or asks about evaluating this task. Reports F1 Score.
Evaluates the ability of machine learning models to distinguish between reference stars and circumstellar exoplanetary disks in high-contrast polarimetric imaging data. It probes representation learning quality through downstream supervised classification and unsupervised clustering tasks. Use when the user wants to benchmark on POLARIS, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates a model's ability to detect online polarization in social media text by classifying statements as polarized or non-polarized. It specifically probes the model's capacity to produce interpretable, structured reasoning alongside binary predictions while handling class imbalance and reducing false negatives. Use when the user wants to benchmark on POLAR @ SemEval-2026, or asks about evaluating this task. Reports macro-F1.
PokeGym evaluates vision-language models' ability to perform long-horizon planning and spatial reasoning in a complex 3D open-world game using only raw RGB observations. It specifically probes visual grounding, autonomous goal decomposition, and physical deadlock recovery, revealing whether models can navigate cluttered environments, interact with objects, and recover from entrapment without explicit state feedback. Use when the user wants to benchmark on PokeGym, or asks about evaluating thi...
Evaluates the ability of Graph Neural Networks to perform node classification while mitigating bias related to a protected attribute (Region). It probes the trade-off between predictive accuracy and group fairness across different GNN architectures. Use when the user wants to benchmark on Pokec-n, or asks about evaluating this task. Reports F1 score.
Evaluates the ability to detect cyber attack campaigns by aligning threat intelligence query graphs with system provenance graphs derived from kernel audit logs. It probes structural pattern matching, causal dependency reasoning, and robustness against malware mutations and benign system noise. Use when the user wants to benchmark on DARPA TC Dataset, Public Malware Reports, or asks about evaluating this task. Reports alignment score.
Evaluates multimodal large language models on fine-grained image understanding and long-form video comprehension tasks. It measures the trade-off between visual token compression efficiency and task accuracy across diverse benchmarks. Use when the user wants to benchmark on MVBench, Video-MME, MLVU, LongVideoBench, MMBench, MMMU_val, or asks about evaluating this task. Reports accuracy.
Evaluates a hybrid graph attention and 3D point cloud neural network's ability to predict quantum chemical properties and molecular physicochemical traits. It probes the model's capacity to integrate 2D topological graph features with 3D spatial geometry for accurate regression and classification of molecular energies and properties. Use when the user wants to benchmark on MoleculeNet, C10, or asks about evaluating this task. Reports MAE, R².
Evaluates vision-language models' embodied reasoning and visual grounding capabilities across three hierarchical stages: referred-object localization, task-driven pointing, and multi-step visual trace prediction in real-world scenarios. Use when the user wants to benchmark on Point-It-Out (PIO), or asks about evaluating this task. Reports score.
Evaluates anomaly detection performance by granting full credit for all points in an anomalous segment if at least one point is detected, often inflating scores for algorithms that merely hit a segment once. Use when the user has predictions and gold and needs to compute point-adjust F1.