Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,817–6,840 of 20,827 skills
Evaluates image classification performance on downsampled variants of ImageNet to test whether lower-resolution datasets can serve as reliable proxies for full-resolution ImageNet in hyperparameter tuning and architecture search. It probes the stability of optimal hyperparameters and model performance across different spatial resolutions while maintaining the original dataset's class structure and image count. Use when the user wants to benchmark on ImageNet32x32, ImageNet64x64, ImageNet16x16...
Evaluates image classification models' accuracy on standard and out-of-distribution datasets. It specifically probes the impact of spatial zooming and foreground/background signal separation on model performance, revealing how much background cues contribute to classification accuracy. Use when the user wants to benchmark on ImageNet, ImageNet-A, ObjectNet, or asks about evaluating this task. Reports top-1 accuracy.
Evaluates the transferability of adversarially robust ImageNet pretraining to downstream classification tasks. It probes whether robustness induces more generalizable and discriminative feature representations compared to standard training, measured under both fixed-feature and full-network fine-tuning settings. Use when the user wants to benchmark on Birdsnap, Caltech-101, Caltech-256, CIFAR-10, CIFAR-100, Describable Textures (DTD), FGVC Aircraft, Food-101, Oxford 102 Flowers, Oxford-IIIT P...
Evaluates the top-1 classification accuracy of a model on the ImageNet dataset. It probes the model's ability to correctly classify images into one of 1000 categories under various training conditions, specifically testing the impact of large minibatch sizes and learning rate scaling strategies on optimization and generalization. Use when the user wants to benchmark on ImageNet, or asks about evaluating this task. Reports top-1 error (%).
Evaluates out-of-distribution (OOD) detection capabilities by measuring how well models assign low confidence to images of objects that do not exist in their training distribution. It probes whether models can reliably distinguish in-distribution classes from novel anomalies without relying on spurious cues. Use when the user wants to benchmark on ImageNet-O, or asks about evaluating this task. Reports AUPR.
Evaluates the trade-off between inference latency, energy consumption, and classification accuracy when dynamically routing image inputs between a lightweight mobile model and a powerful cloud model using a learned neural multiplexer. Use when the user wants to benchmark on ImageNet ILSVRC 2012, or asks about evaluating this task. Reports accuracy.
Evaluates class-conditional image generation fidelity and diversity on ImageNet 256x256. It measures how closely the distribution of generated images matches real images and how well the model covers all classes. Use when the user wants to benchmark on ImageNet, or asks about evaluating this task. Reports FID.
Evaluates text-to-image generation models on photorealism, image-text alignment, and compositional reasoning using standard dataset metrics and human preference studies. Use when the user wants to benchmark on MS-COCO, DrawBench, or asks about evaluating this task. Reports FID-30K.
Evaluates vision-language models' ability to extract structural code (HTML, LaTeX, LilyPond) from images. It uses a round-trip validation pipeline where generated code is rendered back to an image and compared to the original using automated similarity metrics. Use when the user wants to benchmark on Image2Struct, or asks about evaluating this task. Reports EMS.
Evaluates the capability of generative models to produce symbolic music (ABC notation) that aligns with a given input image. It probes both the intrinsic musical quality of the generated output and the semantic/emotional consistency between the source image and the resulting composition. Use when the user wants to benchmark on Image-to-Music test set [[30]], or asks about evaluating this task. Reports Music Quality Level.
Evaluates a model's ability to retrieve relevant images given a text query and vice versa. It probes cross-modal alignment and ranking capabilities under both standard test-set and large-scale candidate-pool settings. Use when the user wants to benchmark on Flickr30k, COCO, or asks about evaluating this task. Reports Recall@K (R@K).
Evaluates the visual fidelity and text-image alignment of generated images. It measures realism and distribution matching using FID, and semantic alignment using CLIP scores. Use when the user wants to benchmark on COCO-2014, or asks about evaluating this task. Reports FID (CLIP features).
Evaluates a model's ability to predict human preferences for text-to-image generation by ranking pairs of images generated from the same text prompt. It measures alignment with human judgment on coherence, fidelity, and aesthetic quality. Use when the user wants to benchmark on ImageReward Test Set, or asks about evaluating this task. Reports Preference Accuracy.
Evaluates the ability of generative models with discrete latents to restore clean images from noisy inputs using a zero-shot, patch-based variational optimization framework. Use when the user wants to benchmark on Standard denoising benchmarks (e.g., House image), or asks about evaluating this task. Reports PSNR.
Evaluates the ability of vision models to learn transferable visual representations and perform accurate image classification across varying data scales and domain shifts. It probes how well patch-based self-attention architectures generalize from large-scale pre-training to standard and low-data downstream recognition tasks. Use when the user wants to benchmark on ImageNet (ILSVRC-2012), or asks about evaluating this task. Reports accuracy.
Evaluates multimodal conversational models on their ability to generate or retrieve engaging, style-conditioned responses grounded in images and dialogue history. It probes retrieval accuracy, generation quality, and human-perceived engagement in multi-turn image-grounded conversations. Use when the user wants to benchmark on IMAGE-CHAT, or asks about evaluating this task. Reports R@1.
Evaluates vision-language models on image captioning and image-text retrieval tasks to measure zero-shot and fine-tuned generalization on long-tail visual concepts and out-of-domain data. Use when the user wants to benchmark on nocaps, COCO Captions, Flickr30K, LocNar Flickr30K, or asks about evaluating this task. Reports CIDEr.
Evaluates a model's ability to generate coherent natural language descriptions for images and align specific image regions with corresponding text segments. It measures both retrieval quality and generation fidelity against human-written references. Use when the user wants to benchmark on Flickr8K, Flickr30K, MSCOCO, or asks about evaluating this task. Reports BLEU.
Evaluates an agent's ability to perform in-context compositional reasoning from image prompts by generalizing learned primitive relations to unseen source-target pairs and complex composite tasks. Use when the user wants to benchmark on 3D Shapes, BitMoji Faces, CLEVR Objects, or asks about evaluating this task. Reports MSE.
Evaluates industrial image anomaly detection algorithms across seven manufacturing datasets under unsupervised, few-shot, continual, and fully supervised settings. It probes both image-level classification and pixel-level localization capabilities, while also measuring computational efficiency like inference speed and GPU memory. Use when the user wants to benchmark on MVTec AD, MVTec LOCO-AD, MPDD, BTAD, MTD, VisA, DAGM, or asks about evaluating this task. Reports Image AUC.
Evaluates large-scale visual recognition capabilities across three core tasks: image classification, single-object localization, and object detection. It probes a model's ability to categorize, localize, and detect objects across 1,000 diverse categories using a dataset of approximately 1 million images. Use when the user wants to benchmark on ILSVRC, or asks about evaluating this task. Reports classification error.
Evaluates the predictive accuracy and learning efficiency of various Inductive Logic Programming (ILP) systems across synthetic grid-world tasks, scalability tests, and standard logical reasoning benchmarks. The protocol measures how well each system generalizes from positive and negative examples to learn correct logic programs under varying domain sizes and example counts. Use when the user wants to benchmark on Robot, Robot2, Member, Benchmark ILP Problems, or asks about evaluating this ta...
Evaluates multimodal models' ability to detect visual illusions (pareidolia) in images, comparing performance across raw, illusory, and low-pass filtered versions. It also measures zero-shot and fine-tuned OCR capabilities on text-containing illusion images to assess perceptual robustness and text recognition under distortion. Use when the user wants to benchmark on IllusionMNIST, IllusionFashionMNIST, IllusionAnimals, IllusionChar, or asks about evaluating this task. Reports Accuracy.
Evaluates instance-level image retrieval capability, measuring a model's ability to correctly rank specific object instances within a massive, domain-diverse image corpus. It probes robustness to background clutter, scale variations, and the effectiveness of global versus local descriptors for re-ranking. Use when the user wants to benchmark on ILIAS, or asks about evaluating this task. Reports mAP@1k.