Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,345–7,368 of 20,845 skills
Evaluates the downstream transfer capability of fractal-based pre-trained visual representations on fine-grained image classification and medical image segmentation tasks. It measures how effectively synthetic Iterated Function System (IFS) pre-training captures transferable features compared to training from scratch or using ImageNet/FractalDB pre-training. Use when the user wants to benchmark on CUB-2011, Stanford Cars, Stanford Dogs, FGVC Aircraft, CIFAR-100, GlaS, or asks about evaluating...
Evaluates trade-offs between inference latency and hardware resource utilization when deploying variational autoencoders on FPGAs using different synthesis frameworks (SNL vs. hls4ml) and quantization levels. Use when the user has predictions and gold and needs to compute Latency.
Evaluates the end-to-end system performance of an FPGA-accelerated machine learning inference service, specifically measuring inference latency and throughput under varying network conditions and concurrent workloads. Use when the user has predictions and gold and needs to compute round-trip inference latency.
Evaluates the end-to-end inference and training latency, power consumption, and energy efficiency of a tensorized neural network hardware accelerator on an FPGA platform, comparing against baseline and prior FPGA accelerators. Use when the user has predictions and gold and needs to compute Latency (ms), Energy Efficiency (GOPS/W).
Evaluates offline-to-online finetuning of model-based reinforcement learning world models on continuous control and visuomotor tasks. It probes the model's ability to adapt to seen and unseen task variations with limited online interactions while mitigating extrapolation errors via uncertainty regularization. Use when the user wants to benchmark on D4RL, xArm, Quadruped Locomotion, Real xArm, or asks about evaluating this task. Reports Success rate (%).
Evaluates how well coordinate-based MLPs with different input feature mappings (none, basic, positional encoding, Gaussian random Fourier features) can learn high-frequency functions across various low-dimensional regression tasks in computer vision and graphics. Use when the user wants to benchmark on Natural images, Text images, 3D shape, Shepp-Logan phantoms, ATLAS dataset, NeRF ATLAS scene, or asks about evaluating this task. Reports PSNR, IoU.
Evaluates the capability of a Fourier-domain low-rank adapter to generate high-quality, diverse images for style transfer and concept editing. It also assesses the adapter's performance on standard language understanding benchmarks compared to baseline adapters like LoRA. Use when the user wants to benchmark on Paintings, Blue-Fire, 3D, Origami, GLUE, or asks about evaluating this task. Reports HPSv2.1.
Evaluates a lightweight foundational model for ECG-based cardiac analysis, specifically testing its ability to classify signals as Normal/Abnormal and perform fine-grained disease classification across multiple cardiac conditions. Use when the user wants to benchmark on PTB-XL, CinC 2017, MedalCare-XL, PTB, or asks about evaluating this task. Reports F1-score.
Evaluates cross-dataset generalization and geographic domain shift in building change detection models. It measures how well models trained on one geographic region or dataset perform when tested on entirely different datasets, highlighting the impact of training data diversity on remote sensing model transferability. Use when the user wants to benchmark on FOTBCD-Binary, LEVIR-CD+, WHU-CD, or asks about evaluating this task. Reports IoU.
Evaluates LLM safeguard robustness against national security and public safety (NSPS) risks by measuring both the model's tendency to generate harmful content in response to adversarial prompts and its tendency to incorrectly refuse legitimate benign requests. Use when the user wants to benchmark on FORTRESS, or asks about evaluating this task. Reports Average Risk Score (ARS).
Evaluates the ability of PDF document parsers to accurately extract mathematical formulas and preserve their semantic meaning. It probes format variability handling, representational non-uniqueness, and semantic equivalence recognition beyond simple character matching. Use when the user wants to benchmark on PDF Formula Extraction Benchmark, or asks about evaluating this task. Reports LLM-as-a-Judge.
Evaluates large language models and speech systems on three endangered Formosan Austronesian languages (Atayal, Amis, Paiwan) across machine translation, automatic speech recognition, and text summarization. It probes zero-shot, few-shot (10-shot), and fine-tuning adaptation capabilities in typologically complex, low-resource settings. Use when the user wants to benchmark on FormosanBench, or asks about evaluating this task. Reports BLEU.
Evaluates large language models' domain knowledge in petroleum geoscience using a 505-question multiple-choice benchmark. It probes understanding across seven specialized subdomains, including petrophysics, reservoir engineering, and drilling, while measuring performance variance by model size, cost, and question difficulty. Use when the user wants to benchmark on FormationEval, or asks about evaluating this task. Reports Accuracy.
Evaluates the ability of auxiliary-task learning methods to mitigate negative transfer and improve target task performance across multi-task, multi-domain, and semi-supervised learning settings. It probes how well a model can dynamically combine or select auxiliary tasks without degrading the primary task. Use when the user wants to benchmark on NYUv2, DomainNet, AliExpress, CIFAR-10, SVHN, or asks about evaluating this task. Reports Δm.
This benchmark probes the ability of multimodal large language models to recognize regular and irregular geometric shapes from images and accurately count their sides. It further evaluates multi-step visual-mathematical reasoning by requiring models to identify multiple shapes, map them to side counts, and compute their sum. Use when the user wants to benchmark on Forgotten Polygons, or asks about evaluating this task. Reports accuracy.
Evaluates pixel-level localization accuracy for detecting AI-generated and traditionally tampered image forgeries. It probes a model's ability to distinguish manipulated regions from authentic content by measuring spatial overlap and detection trade-offs against ground-truth masks. Use when the user wants to benchmark on OpenSDID, GIT10K, CocoGlide, Inpaint32K, IMD2020, NIST16, CASIA, or asks about evaluating this task. Reports F1-score.
Evaluates multimodal large language models on fine-grained manufacturing tasks, including workpiece verification, surface defect inspection, and assembly verification. It probes the models' ability to combine visual grounding with domain-specific knowledge to identify anomalies or classify conditions. Use when the user wants to benchmark on FORGE, or asks about evaluating this task. Reports exact-match accuracy.
Evaluates a vision-language agent's ability to perform joint change detection and natural language captioning on bi-temporal remote sensing imagery. It probes the model's capacity to segment deforestation and built-environment changes at the pixel level while generating accurate semantic descriptions of those changes. Use when the user wants to benchmark on Forest-Change, LEVIR-MCI-Trees, or asks about evaluating this task. Reports MIoU.
Evaluates a fine-tuned LLM's ability to predict future biomedical concepts and clinical disorders from patient clinical timelines. It measures concept prediction accuracy via precision and recall across different temporal windows and candidate counts, and assesses clinical risk forecasting by checking how many of the top-5 predicted disorders match the ground truth for the next month. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports Precision.
Evaluates a model's ability to perform temporally grounded, multimodal video understanding in surveillance settings. It probes precise temporal localization, identity-based search, and complex reasoning over long videos using text-only or image+text queries. Use when the user wants to benchmark on ForeSeaQA, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates Large Vision Language Models (LVLMs) on their ability to detect and attribute image or video forgeries. It probes generalization and reasoning capabilities across five dimensions: semantics, modalities, tasks, forgery types, and generation models, using multi-choice visual questions. Use when the user wants to benchmark on Forensics-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates hierarchical multi-label classification of academic papers into a taxonomy of 170 research fields. It tests zero-shot/few-shot prompting and weakly-labeled data integration for field-of-research prediction. Use when the user wants to benchmark on FoRC4CL 2025, or asks about evaluating this task. Reports Micro-F1.
Evaluates the quality of universal image embeddings for flat object retrieval across diverse 2D domains (e.g., logos, paintings, currency) under varying visual distortions. It probes both candidate rank accuracy and the matching score margin to assess out-of-distribution generalization. Use when the user wants to benchmark on FORB, or asks about evaluating this task. Reports mAP@5.
This benchmark evaluates the robustness and accuracy of Text-to-SQL systems when translating natural language questions into SQL queries across different database schema designs. It probes how data model complexity, training data size, and language model scale impact execution accuracy on real-world user queries. Use when the user wants to benchmark on FootballDB, or asks about evaluating this task. Reports exact execution matching (EX).