Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,961–6,984 of 20,831 skills
Evaluates an embodied AI agent's ability to navigate indoor 3D environments to find specific object categories using RGB-D observations. It measures both navigation quality (success and path efficiency) and computational efficiency (latency, memory, and skip ratio) on a large-scale dataset. Use when the user wants to benchmark on HabitatMatterport3D (HM3D), or asks about evaluating this task. Reports SPL.
Evaluates the quality of self-supervised visual representations learned from histopathology images by measuring downstream classification performance at patch, slide, and patient levels. It probes the model's ability to capture hierarchical pathological structures and align them with clinical diagnostic categories. Use when the user wants to benchmark on OpenSRH, TCGA, or asks about evaluating this task. Reports kNN classification accuracy (ACC).
Evaluates large language models' multilingual comprehension of Hong Kong-specific knowledge, Cantonese linguistic capabilities, and reasoning across STEM, social sciences, and humanities in both Traditional and Simplified Chinese. Use when the user wants to benchmark on HKMMLU, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of LLM-based sequential recommendation models to predict the next item in a user's interaction history. It specifically probes how well models capture temporal dynamics by incorporating irregular time intervals between interactions, and assesses performance under warm and cold-start conditions. Use when the user wants to benchmark on Amazon Reviews (Video Games, CDs and Vinyl, Books), or asks about evaluating this task. Reports Hit Ratio@1.
This benchmark evaluates graph-based generative models for their ability to produce biologically relevant, drug-like hit molecules rather than just chemically valid structures. It probes the models' capacity to satisfy strict medicinal chemistry constraints, maintain distributional similarity to known bioactive compounds, and achieve strong predicted binding affinity to specific protein targets. Use when the user wants to benchmark on REINVENT Dataset, Hit-like Dataset, Target-Specific Ligand...
Multi-class histopathological image classification for cancer diagnosis across four tissue types (breast, prostate, bone, cervical). It probes the model's ability to extract robust morphological features from stained whole-slide image tiles without data augmentation. Use when the user wants to benchmark on ICIAR2018, SIPAkMeD, SICAPv2, UT-Osteosarcoma, or asks about evaluating this task. Reports accuracy.
Evaluates a unified GAN framework's ability to perform stain-invariant segmentation of glomeruli in renal histopathology. It tests generalization across multiple known staining modalities and unseen stainings, measuring how well the model maintains segmentation accuracy despite domain shifts in histological appearance. Use when the user wants to benchmark on AIDPATH & Custom PAS dataset, or asks about evaluating this task. Reports F1.
Evaluates LLMs' ability to accurately transcribe historical 18th-century Russian documents while preserving period-specific orthography and avoiding anachronistic character insertions. It probes both standard OCR accuracy and historical fidelity under varying input contexts and prompt strategies. Use when the user wants to benchmark on 18th-century Russian Civil Font Texts, or asks about evaluating this task. Reports CER.
Evaluates deep learning models on binary classification of histopathology tissue tiles as malignant or benign. It probes the capacity of multi-stream architectures to capture diverse morphological and textural features for medical image grading. Use when the user wants to benchmark on CAMELYON16, Invasive Ductal Carcinoma (IDC), or asks about evaluating this task. Reports accuracy.
Evaluates graph neural network explainers on histopathology images by measuring how well they identify critical tumor nuclei and preserve model fidelity. It probes the explainer's ability to extract global, class-specific patterns and produce accurate instance-level importance maps for downstream nuclei classification. Use when the user wants to benchmark on BRACS, BACH, BreCaHAD, CRC, or asks about evaluating this task. Reports macro-averaged F1 score.
Evaluates a model's ability to generalize to out-of-distribution domains (different hospitals or staining protocols) in histopathology image classification. It measures classification accuracy on held-out OOD validation and test splits, alongside the reconstruction quality of self-supervised generative augmentation. Use when the user wants to benchmark on CAMELYON17-WILDS, Epithelium-Stroma, or asks about evaluating this task. Reports Accuracy (%).
Evaluates the robustness of vision-language models (VLMs) and test-time adaptation (TTA) methods when applied to histopathology images under realistic domain shifts. It probes how well models maintain classification accuracy when exposed to synthetic corruptions like staining variations, dust, blurring, and noise that mimic real-world clinical imaging artifacts. Use when the user wants to benchmark on NCT-7K, NCT-100K, LC25000, SkinCancer, RenalCell, MHIST, or asks about evaluating this task....
Evaluates the linguistic quality and cultural appropriateness of the Histoires Morales dataset. It measures reference-free translation accuracy and assesses whether moral norms and actions align with native French speakers' cultural values. Use when the user wants to benchmark on Histoires Morales, or asks about evaluating this task. Reports CometKiwi22.
Evaluates the prognostic and molecular predictive value of 38 automated histomic features extracted from H&E whole-slide images across 21 solid-tumor cancer types. It probes whether purely morphological patterns can recover canonical biology, predict survival outcomes, and correlate with gene expression, pathway activity, and immune subtypes. Use when the user wants to benchmark on TCGA Pan-Cancer H&E Cohort, or asks about evaluating this task. Reports Cox proportional-hazards model.
Evaluates vision-language foundation models on diverse histopathology clinical tasks (detection, subtyping, grading, mutation prediction) to assess their robustness to textual/visual perturbations, magnification changes, stain normalization, and model calibration. Use when the user wants to benchmark on CRC-100K, BreakHist, DataBiox, GasHisSDB, Breast IDC, LC25000-lung, or asks about evaluating this task. Reports balanced accuracy.
Evaluates the ability of language models to recognize and classify named entities (PERSON, ORGANIZATION, LOCATION, PRODUCT, DATE) in historical Romanian newspaper texts across four distinct geographical regions. Use when the user wants to benchmark on HistNERo, or asks about evaluating this task. Reports strict F1-score.
Evaluates a model's ability to generate clinical histopathology reports from gigapixel whole slide images (WSIs). It probes cross-modal alignment between dense visual patches and concise textual descriptions using standard natural language generation metrics. Use when the user wants to benchmark on TCGA WSI-Report, or asks about evaluating this task. Reports BLEU-4.
Evaluates a vision-language model's ability to generate accurate and clinically relevant histopathology reports from whole slide images (WSIs). It probes lexical overlap, semantic coherence, and medical entity coverage in generated text. Use when the user wants to benchmark on HistGen, or asks about evaluating this task. Reports BLEU-4.
Evaluates large language models across five hierarchical cognitive stages of scientific inquiry, ranging from foundational factual recall and literature parsing to advanced synthesis, literature review generation, and data-driven scientific discovery. It probes multimodal comprehension, cross-lingual reasoning, and computational problem-solving across six scientific disciplines. Use when the user wants to benchmark on HiSciBench, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to perform 3D human-in-scene multimodal understanding through open-ended question answering. It probes capabilities in activity recognition, spatial relationship reasoning, and human-object interaction analysis within dynamic 3D environments. Use when the user wants to benchmark on HIS-Bench, or asks about evaluating this task. Reports HIS-Bench score.
Evaluates the answer quality of Retrieval-Augmented Generation (RAG) systems across specialized domains. It measures how well generated responses address queries in terms of comprehensiveness, empowerment, diversity, and overall performance using pairwise LLM-as-a-judge comparisons. Use when the user wants to benchmark on UltraDomain, or asks about evaluating this task. Reports win rate.
Evaluates histopathology image retrieval and classification performance using high-order texture features (Gram barcodes) extracted from CNN layers. It probes the model's ability to capture tissue texture patterns for accurate image matching and class prediction. Use when the user wants to benchmark on KimiaPath24, CRC, EMC, or asks about evaluating this task. Reports η_total, Accuracy.
Evaluates multimodal agents' ability to reason over personalized, device-scale file systems. It probes long-horizon cross-file retrieval, multimodal perception, and evidence-grounded factual retention under strict profile-isolation constraints. Use when the user wants to benchmark on HippoCamp, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates multimodal physical reasoning and advanced problem-solving capabilities on international and regional physics Olympiad exams. It probes a model's ability to interpret complex diagrams, data, and text, perform multi-step logical derivations, and produce accurate solutions under strict, official scoring rubrics. Use when the user wants to benchmark on HiPhO, or asks about evaluating this task. Reports exam score.