Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,529–6,552 of 20,836 skills
Evaluates the ability of state-space models (Mamba/SSD-Mamba) and transformers to perform statutory classification and case law retrieval on long-context legal documents. It probes how well models capture fine-grained semantic distinctions and maintain global coherence over thousands of tokens while balancing accuracy with computational throughput. Use when the user wants to benchmark on SCOTUS, ILDC, ECtHR, EUR-Lex, or asks about evaluating this task. Reports Accuracy.
Evaluates code generation and algorithmic reasoning capabilities on competitive programming problems. It specifically probes temporal robustness by testing on problems released after a strict cutoff date to detect data contamination, and measures performance across difficulty levels and algorithmic topics. Use when the user wants to benchmark on LeetCodeDataset, or asks about evaluating this task. Reports pass@1.
Probes a model's ability to retrieve relevant Chinese criminal case documents from a large corpus based on legal queries. It specifically tests alignment with multi-dimensional legal relevance criteria, including case characterization, penalty matching, and procedural similarity. Use when the user wants to benchmark on LeCaRDv2, or asks about evaluating this task. Reports Recall@K.
Evaluates end-to-end learned image signal processing (ISP) pipelines that map mobile RAW sensor data to high-fidelity RGB images. It probes the trade-off between image reconstruction fidelity, subjective visual quality, and real-time inference efficiency on mobile hardware. Use when the user wants to benchmark on Fujifilm UltraISP dataset, or asks about evaluating this task. Reports PSNR.
Evaluates a unified framework for multi-task domain adaptation few-shot learning across image classification, object detection, and video classification. It probes the model's ability to adapt to new domains and scale label budgets incrementally from 1-shot to full dataset size. Use when the user wants to benchmark on DomainNet, Office-Home, Office31, Pool and Car, xView, UCF101, or asks about evaluating this task. Reports accuracy.
Evaluates the quality of answers generated by RAG systems across specialized domains. It probes the model's ability to retrieve relevant information, synthesize comprehensive responses, and maintain diversity and practical utility. Additionally, it measures retrieval efficiency and the impact of structural knowledge on generation. Use when the user wants to benchmark on UltraDomain, or asks about evaluating this task. Reports Comprehensiveness.
Evaluates a language model's ability to generate correct, step-by-step formal proof tactics for mathematical statements within the Lean 4 proof assistant. Use when the user wants to benchmark on miniF2F, or asks about evaluating this task. Reports solve_rate.
Evaluates an LLM-based framework's ability to dynamically generate research leaderboards by measuring topic relevance, content quality (coverage, recency, structure), and generation speed compared to manual curation. Use when the user has predictions and gold and needs to compute Leaderboard Content Quality.
Evaluates vision and vision-language models on plant disease diagnosis, including fine-grained image classification, few-shot adaptation, and zero-shot visual question answering. It probes the models' ability to recognize subtle visual symptoms, reason over taxonomic pathogen information, and generalize across agricultural domains. Use when the user wants to benchmark on LeafNet, or asks about evaluating this task. Reports Accuracy.
Evaluates federated learning algorithms under realistic constraints including device-level data skew, heterogeneous data distributions, and communication bottlenecks. It measures model accuracy after federated training across multiple simulated devices. Use when the user wants to benchmark on Shakespeare, Sent140, FEMNIST, CelebA, Synthetic, Reddit, or asks about evaluating this task. Reports AccuracyTop1.
Evaluates whether pre-trained Recognizing Textual Entailment (RTE) models can generalize to unseen task-dataset-metric (TDM) extraction pairs in a zero-shot setting. It probes whether models learn genuine semantic entailment or merely memorize training distribution patterns. Use when the user wants to benchmark on LEADERBOARDS, or asks about evaluating this task. Reports macro F1.
Evaluates a model's ability to verify whether a candidate (Task, Dataset, Metric) triple is actually used or mentioned in a specific AI research paper. The task is framed as a natural language inference problem where the model must distinguish between valid triples and randomly sampled invalid ones. Use when the user wants to benchmark on AI Research Paper Collection, or asks about evaluating this task. Reports micro-F1.
Evaluates the closed-loop driving performance and generalization of end-to-end autonomous driving policies in simulation and on real-world datasets. It probes the model's ability to navigate long-horizon routes, handle diverse weather and lighting conditions, and transfer synthetic pre-training to real-world driving scenarios without violating traffic rules. Use when the user wants to benchmark on CARLA Town13, Bench2Drive, Longest6 v2, NAVSIM v1, NAVSIM v2, WOD-E2E, or asks about evaluating ...
Evaluates the inference latency, throughput, and SLA compliance of a dynamic batching system under varying request arrival rates and diverse DNN workloads. Use when the user wants to benchmark on ResNet, GNMT, Transformer, VGGNet, MobileNet, LAS, BERT, or asks about evaluating this task. Reports SLA violation rate.
Evaluates layout-guided image generation models on their ability to follow spatial control instructions (number, position, size, shape) across in-distribution and out-of-distribution layouts. Probes generalization to arbitrary object configurations and fine-grained spatial reasoning. Use when the user wants to benchmark on CLEVR, LayoutBench, or asks about evaluating this task. Reports AP (AP50).
Evaluates a model's ability to generate high-fidelity images conditioned on spatial layouts and text descriptions, measuring both perceptual quality and precise object-level spatial alignment. Use when the user wants to benchmark on COCO-3K, HiCo-7K, or asks about evaluating this task. Reports FID.
Evaluates optical flow estimation on non-Lambertian surfaces (transparent, reflective, diffuse) and multi-layer scenes. It probes a model's ability to predict flow through transparent occluders and handle complex material properties without relying on test-time optimizations. Use when the user wants to benchmark on LayeredFlow, or asks about evaluating this task. Reports EPE.
Evaluates a model's ability to decompose raster graphic designs into a sequence of re-editable layers. It measures visual reconstruction quality and the number of edits required to match a ground-truth layer structure, accounting for the ill-posed nature of layer ordering. Use when the user wants to benchmark on Crello, or asks about evaluating this task. Reports RGB L1, Alpha IoU.
This benchmark evaluates the capability of models to detect and temporally localize content-driven audio-visual forgeries in long videos. It probes multimodal boundary matching and temporal manipulation detection by requiring models to identify fake segments and predict their precise start and end timestamps. Use when the user wants to benchmark on LAV-DF, or asks about evaluating this task. Reports AP@0.5.
Evaluates Latvian-specific encoder models on lightweight diagnostic tasks, morphosyntactic parsing, and semantic representation quality to benchmark low-resource language modeling capabilities. Use when the user wants to benchmark on EuroEval Latvian diagnostics, COPA (Latvian), Universal Dependencies Latvian treebank (UD v2.16), Latvian WSD dataset, or asks about evaluating this task. Reports MCC.
Evaluates a model's ability to detect unanswerable Text-to-SQL queries by analyzing intermediate hidden activations, aiming to prevent hallucinated SQL generation and unsafe execution. It probes whether the system can reliably distinguish between answerable and unanswerable prompts across diverse domains and linguistic ambiguities. Use when the user wants to benchmark on TriageSQL, AMBROSIA, SQuAD 2.0, MD-Enterprise, or asks about evaluating this task. Reports F1.
Evaluates a language model's reasoning, coding, and general knowledge capabilities using a suite of standard academic benchmarks. It specifically probes how test-time compute scaling (via recurrent depth) impacts performance across mathematical, coding, and commonsense reasoning tasks. Use when the user wants to benchmark on GSM8K, MATH (Minerva), MathQA, MBPP, HumanEval, ARC-E, ARC-C, HellaSwag, MMLU, OBQA, PiQA, SciQ, WinoGrande, or asks about evaluating this task. Reports flexible extract ...
Measures the inference latency of binarized, 8-bit, and 32-bit convolutional layers on edge devices to evaluate the efficiency and speedup of the Larq Compute Engine framework compared to standard implementations. Use when the user has predictions and gold and needs to compute latency.
Evaluates the real-time performance and observability accuracy of an eBPF-based tracing library by measuring request throughput and tail latency under inference workloads. It verifies the framework's ability to disambiguate request boundaries from streaming system calls without application instrumentation. Use when the user has predictions and gold and needs to compute latency statistics.