Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,177–7,200 of 20,839 skills
Evaluates the effectiveness of LLM-based prompt optimizers across complex reasoning, knowledge-intensive, and common NLP tasks. It measures how much optimized prompts improve model performance compared to baseline prompts and other optimization methods. Use when the user wants to benchmark on Big-Bench Hard (BBH), GSM8K, MMLU, WSC, WebNLG, or asks about evaluating this task. Reports average accuracy.
Evaluates tracking performance on a large-scale dataset of 10,000 videos with diverse object categories, testing generalization to real-world scenarios with language guidance. Use when the user wants to benchmark on GOT-10k, or asks about evaluating this task. Reports AO.
Evaluates an LLM's ability to generate correct API invocation code from natural language prompts, with or without retrieved documentation. It measures how well the model selects the appropriate API, avoids hallucinating non-existent APIs, and respects functional constraints like accuracy thresholds. Use when the user wants to benchmark on APIBench, or asks about evaluating this task. Reports AST accuracy.
Evaluates reinforcement learning agents' ability to learn multi-agent coordination, strategic decision-making, and long-horizon planning in a physics-based 3D football simulation. It probes sample efficiency, reward shaping robustness, and performance under varying opponent difficulty levels. Use when the user wants to benchmark on Google Research Football (Football Benchmarks), or asks about evaluating this task. Reports average goal difference.
This benchmark evaluates the capability of large language models to perform a wide range of financial natural language processing tasks in both English and Chinese. It probes domain-specific understanding, information extraction, reasoning, and generation across sentiment analysis, classification, entity/relation extraction, summarization, question answering, and stock movement prediction. Use when the user wants to benchmark on FPB, Fiqa-SA, Headlines, FOMC, lendingclub, NER, FinRE, CFA, EDT...
Evaluates cross-domain generalization of emotion classification models by measuring how well a model fine-tuned on GoEmotions transfers to external benchmarks with limited labeled data, compared to training from scratch. Use when the user wants to benchmark on ISEAR, EmoInt, Emotion-Stimulus, or asks about evaluating this task. Reports average F1-score.
Evaluates large multimodal models' ability to perform goal-oriented embodied navigation in complex urban 3D airspace. It probes geometric perception, cross-view understanding, spatial imagination, and long-term memory by requiring models to navigate from a start point to a semantic goal using visual observations and historical context. Use when the user wants to benchmark on Embodied Navigation Benchmark, or asks about evaluating this task. Reports SR.
Evaluates a self-supervised, classification-based anomaly detection framework (GOAD) on its ability to distinguish normal data from anomalies without labeled anomalies during training. It probes the model's robustness to data contamination and adversarial attacks across image and tabular domains. Use when the user wants to benchmark on CIFAR-10, FashionMNIST, Arrhythmia, Thyroid, KDD, KDDRev, or asks about evaluating this task. Reports ROC-AUC.
Evaluates the quality, robustness, and feasibility of perturbation-based GNN explainers. It probes whether generated subgraphs (factual or counterfactual) reliably preserve or flip model predictions, remain stable under topological or architectural perturbations, and satisfy domain-specific structural constraints. Use when the user wants to benchmark on Mutagenicity, Proteins, IMDB-B, AIDS, MUTAG, NCI1, Graph-SST2, DD, REDDIT-B, ogbg-molhiv, Tree-Cycles, Tree-Grid, BA-Shapes, or asks about ev...
Evaluates Graph Neural Networks for classifying emotional states from 32-channel EEG signals. It probes spatial-temporal feature extraction, cross-subject generalization, and robustness to hyperparameter settings in neuroscience signal processing. Use when the user wants to benchmark on FACED, or asks about evaluating this task. Reports Accuracy.
Compares graph neural networks against deep fully-connected feedforward networks for binary classification of b-quark pairs in top-quark-antiquark collisions at the LHC. It probes whether explicit relational inductive biases in GNNs outperform permutation-invariant DNNs when provided with equivalent kinematic and relational features. Use when the user wants to benchmark on LHC t\bar{t} b\bar{b} event classification, or asks about evaluating this task. Reports mean ROC-AUC ($\mu_{\mathrm{AUC}}$).
Evaluates the runtime and memory efficiency of low-level Graph Neural Network and sparse tensor operations on NVIDIA A100 GPUs. It probes how input sparsity, tensor dimensions, and reduce factors affect computational overhead when operations are pushed to near-full GPU memory capacity. Use when the user has predictions and gold and needs to compute median_runtime.
Evaluates Graph Neural Networks and traditional ML models on a compiled Internet routing dataset for link prediction and node classification. It probes the ability to infer missing AS-AS connections and predict voluntary PeeringDB attributes under conditions of graph sparsity and label imbalance. Use when the user wants to benchmark on Internet Routing AS Graph Dataset, or asks about evaluating this task. Reports binary cross entropy.
Evaluates the quality and interpretability of explanations generated by various Graph Neural Network (GNN) explainers across different architectures and graph datasets. It probes how well explanations align with human expectations (plausibility) and model decision logic (fidelity). Use when the user wants to benchmark on Grid, Grid-House, Stars, House-Color, or asks about evaluating this task. Reports F1-Fidelity.
Evaluates an LLM-guided framework's ability to automatically propose and refine Graph Neural Network architectures for node classification across diverse graph datasets, including out-of-distribution and heterophilic graphs, without requiring extensive training or search. Use when the user wants to benchmark on NAS-Bench-Graph & OOD Graphs (Cora, Citeseer, PubMed, CS, Physics, Photo, Computer, ogbn-arXiv, DBLP, Flickr, Actor), or asks about evaluating this task. Reports accuracy.
This evaluation probes how graph neural network (GNN) performance depends on the ratio of observed training nodes to feature dimensionality (N_obs/D). It tests whether standard benchmarks are biased toward low-dimensional regimes and evaluates the effectiveness of decoupling feature extraction from graph propagation. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, or asks about evaluating this task. Reports Accuracy.
Evaluates the expressiveness and practical performance of GNN-AK, a framework that replaces star-shaped neighbor aggregation with subgraph-based encoding in Message Passing Neural Networks. It probes the model's ability to distinguish complex graph structures (e.g., strongly regular graphs, substructures) and predict graph-level properties on standard benchmarks. Use when the user wants to benchmark on ZINC-12K, CIFAR10, PATTERN, MolHIV, MolPCBA, EXP, SR25, or asks about evaluating this task....
Evaluates the adversarial robustness of Graph Neural Networks against node/edge injection and modification attacks, comparing Hamiltonian-based models against standard GNNs and defense baselines. Use when the user wants to benchmark on Cora, Citeseer, Pubmed, Coauthor, Computers, Ogbn-Arxiv, Polblogs, or asks about evaluating this task. Reports accuracy.
Evaluates the accuracy of density functionals in predicting broad molecular properties, including atomization energies, barrier heights, and noncovalent interactions, relative to high-level quantum chemistry reference data. Use when the user wants to benchmark on GMTKN55 (specifically GMTKN53 subset), or asks about evaluating this task. Reports WTMAD1.
Evaluates a model's ability to perform pixel-level semantic segmentation for drivable areas and road anomalies using multi-modal visual inputs. It specifically probes how effectively networks can fuse RGB imagery with depth-related features (e.g., transformed disparity) to improve detection accuracy for ground mobile robots. Use when the user wants to benchmark on GMRP, KITTI road, KITTI semantic segmentation, or asks about evaluating this task. Reports IoU.
Evaluates a vision-language model's ability to understand and reason over diverse medical imaging modalities (CT, MRI, X-ray, pathology slides) to answer clinical questions, diagnose diseases, and perform anatomical or lesion recognition tasks. Use when the user wants to benchmark on PMCVQA, PathVQA, VQA-RAD, SLAKE, OmniMedVQA, GMAI-MMBench, MMMU Health & Medicine track, or asks about evaluating this task. Reports accuracy.
Tests the ability of language models to automatically classify documents into predefined group membership categories (e.g., gender, geographic location) for group fairness evaluation in information retrieval. Use when the user wants to benchmark on TREC fair ranking track 2021, TREC fair ranking track 2022, NTCIR fairweb1 (Chuweb-21D), or asks about evaluating this task. Reports accuracy.
Evaluates general language understanding across classification, regression, and entailment tasks, alongside language modeling capability. It also measures computational efficiency and internal attention allocation strategies under resource constraints. Use when the user wants to benchmark on GLUE Benchmark, WikiText-103, or asks about evaluating this task. Reports MNLI-m Accuracy.
Evaluates the generalization and downstream performance of pretrained language models on a suite of natural language understanding tasks (GLUE) and reading comprehension (SQuAD 2.0). Use when the user wants to benchmark on GLUE, SQuAD 2.0, or asks about evaluating this task. Reports GLUE.