Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,465–7,488 of 20,843 skills
Evaluates whether LLMs can perform domain-level conceptual understanding and precise financial reasoning across multilingual professional certification exams. It probes analytical rigor and regulatory knowledge integration rather than simple factual recall. Use when the user wants to benchmark on EFPA, GRFinQA, CFA, CPA, BBF, SAHM, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates large language models' ability to perform knowledge-intensive mathematical reasoning within the finance domain. It requires models to integrate college-level financial knowledge with both textual descriptions and tabular data to solve complex problems. Use when the user wants to benchmark on FinanceMATH, or asks about evaluating this task. Reports accuracy.
Evaluates large language models' ability to answer financial questions using various retrieval and context strategies, probing numerical reasoning, factuality, and handling of structured or long documents. Use when the user wants to benchmark on FinanceBench, or asks about evaluating this task. Reports correct answer.
Evaluates large language models on Finnish language capabilities across reading comprehension, commonsense reasoning, sentiment analysis, world knowledge, truthfulness, and alignment. It probes both multiple-choice and generative capabilities under varying prompt formulations (cloze vs. multiple-choice) and shot configurations. Use when the user wants to benchmark on FIN-bench-v2, or asks about evaluating this task. Reports normalized accuracy.
Evaluates Finnish large language models on a curated suite of 11 tasks spanning arithmetic, reasoning, general knowledge, emotion classification, and linguistic understanding. It probes few-shot generalization and task-specific capabilities in a low-resource language setting. Use when the user wants to benchmark on FIN-bench, or asks about evaluating this task. Reports mean accuracy.
Evaluates a model's ability to detect and classify filler words (e.g., 'uh', 'um') in naturalistic speech recordings. It probes temporal localization accuracy and fine-grained acoustic classification under varying evaluation granularities. Use when the user wants to benchmark on PodcastFillers, or asks about evaluating this task. Reports F1.
Evaluates text classification capability in low-resource settings for the Filipino language, specifically measuring model robustness and performance degradation as training data size is systematically reduced. Use when the user wants to benchmark on Hate Speech, Dengue, or asks about evaluating this task. Reports accuracy, hamming loss.
Evaluates AI agents' ability to personalize based on file-system behavioral traces across procedural, semantic, and episodic memory channels. It probes attribute recognition, behavioral inference, anomaly detection, and grounding using simulated, multimodal, and real-world settings. Use when the user wants to benchmark on FileGramBench, or asks about evaluating this task. Reports accuracy.
Evaluates LLMs' ability to understand and generate text in Filipino, Tagalog, and Cebuano across cultural knowledge, classical NLP tasks, reading comprehension, and text generation. It probes cultural alignment, factual recall, linguistic processing, and translation capabilities in low-resource Southeast Asian languages. Use when the user wants to benchmark on FilBench, or asks about evaluating this task. Reports FilBench Score.
Evaluates the ability of text-to-illustration models to generate publication-ready scientific figures that balance structural fidelity, visual aesthetics, and communicative clarity based on long-form scientific text. Use when the user wants to benchmark on FigureBench, or asks about evaluating this task. Reports Overall score.
This benchmark evaluates the ability of vision-language and image-editing models to perform semantically correct, structure-aware modifications to scientific charts. It probes whether models can follow precise editing instructions while preserving data-encoding consistency, axis coherence, and legend integrity, rather than merely producing pixel-level visual similarity. Use when the user wants to benchmark on FigEdit, or asks about evaluating this task. Reports Instruction-following score.
Evaluates language models' ability to interpret nonliteral, creative metaphors by selecting the correct literal meaning from two opposing options, and generating sensible interpretations for novel metaphors. It probes commonsense grounding and contextual understanding beyond literal paraphrase tasks. Use when the user wants to benchmark on Fig-QA, or asks about evaluating this task. Reports accuracy.
Evaluates large language models' ability to follow complex financial instructions, with a strong emphasis on precise adherence to formatting, structural constraints, and conditional styling requirements. It probes whether models can maintain procedural compliance rather than just semantic correctness. Use when the user wants to benchmark on FIFE, or asks about evaluating this task. Reports Strict compliance.
Probes whether input attribution methods accurately reflect token importance for model predictions, particularly when inputs are adversarially perturbed or masked out-of-distribution. It evaluates the consistency of fidelity scores across different model architectures and under various adversarial attacks. Use when the user has predictions and gold and needs to compute fidelity.
Measures the distributional similarity between original GAN-generated images and their semantically manipulated counterparts. It evaluates whether a latent space transformation preserves overall image quality and realism while altering specific attributes. Use when the user has predictions and gold and needs to compute FID.
Evaluates the ability of shallow and deep learning models to predict click-through rates (CTR) on large-scale ad impression datasets. It probes how well architectures can model high-order feature interactions and dynamically weight feature importance using bilinear functions and Squeeze-Excitation networks. Use when the user wants to benchmark on Criteo, Avazu, or asks about evaluating this task. Reports AUC.
Evaluates an LLM verifier's ability to rank multiple candidate answers by their factual correctness and error severity across single-hop and multi-hop questions, using external knowledge retrieval and fine-grained scoring. Use when the user wants to benchmark on FGVeriBench, or asks about evaluating this task. Reports Kendall-tau.
Evaluates a fine-grained selective similarity integration framework for drug-target interaction prediction. It tests the model's ability to dynamically weight multiple drug and target similarity views based on local interaction consistency to predict binary interaction labels. Use when the user wants to benchmark on Nuclear Receptors (NR), G-protein coupled receptors (GPCR), Ion Channel (IC), Enzyme (E), Luo, or asks about evaluating this task. Reports AUC.
Evaluates session-based recommendation models by predicting the next item in a user's clickstream session. It probes the model's ability to capture latent, non-temporal transition patterns and item order within a session graph. Use when the user wants to benchmark on Yoochoose, Diginetica, or asks about evaluating this task. Reports R@20.
Evaluates session-based recommendation models on e-commerce clickstream data by predicting the next item in a session using graph neural networks and cross-session information. Use when the user wants to benchmark on Yoochoose1/64, Yoochoose1/4, Diginetica, or asks about evaluating this task. Reports R@20, MRR@20.
Evaluates fine-grained named entity recognition (FgNER) capabilities across multiple languages by measuring how accurately models identify and classify specific entity types within text sequences. The protocol assesses sequence labeling performance using standard span-based metrics. Use when the user wants to benchmark on Various NER datasets (cited in text), or asks about evaluating this task. Reports F1-score.
Evaluates a functional generative network for medium-range probabilistic weather forecasting against operational ground truth (HRES-fc0) and a diffusion-based baseline (GenCast). Probes the model's ability to capture joint spatial structures and predict tropical cyclone tracks using deterministic and probabilistic scoring rules. Use when the user wants to benchmark on HRES-fc0, ERA5, or asks about evaluating this task. Reports probabilistic metrics.
Tests functional-group reasoning capability by asking models to predict how molecular properties change when specific functional groups are added, removed, or modified at given positions. Use when the user wants to benchmark on FGBench, or asks about evaluating this task. Reports Accuracy (Acc).
Evaluates the transfer learning capability of a feature-guided masked autoencoder on remote sensing imagery. It probes the model's ability to adapt to downstream scene classification and semantic segmentation tasks across multispectral and SAR modalities using linear probing and fine-tuning protocols. Use when the user wants to benchmark on BigEarthNet-MM, BigEarthNet-SAR, EuroSAT, EuroSAT-SAR, DFC2020, or asks about evaluating this task. Reports mAP.