Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,257–5,280 of 23,579 skills
Evaluates whether direct classification of RAW sensor data achieves accuracy comparable to traditional RAW-to-RGB converted images, while measuring computational efficiency gains from skipping the conversion pipeline. Use when the user wants to benchmark on Custom RAW/RGB Dataset, or asks about evaluating this task. Reports top-1 classification accuracy.
Evaluates interpretability methods' ability to disentangle polysemantic language model representations by isolating causal attributes through activation interventions on residual stream features. Use when the user wants to benchmark on RAVEL, or asks about evaluating this task. Reports Disentanglescore.
Evaluates the fidelity of a simulated noisy speech generator against real VHF/UHF transmitted audio, and measures the downstream robustness of automatic speech recognition (ASR) models trained on the simulated data. Use when the user wants to benchmark on RATS Channel A, or asks about evaluating this task. Reports MSSL, WER.
This evaluation probes the robustness of end-to-end automatic speech recognition systems in noisy acoustic conditions. It measures how well a model preserves speech intelligibility and correctly transcribes utterances when background noise is present, specifically testing the mitigation of over-suppression artifacts during joint speech enhancement and recognition. Use when the user wants to benchmark on RATS Channel-A, or asks about evaluating this task. Reports WER(%).
This evaluation probes the robustness of automatic speech recognition (ASR) systems when trained on extremely limited in-domain noisy data. It measures how well a model can generalize to real-world noisy conditions by leveraging synthetic noisy data generated via a GAN, compared to traditional data augmentation and fine-tuning baselines. Use when the user wants to benchmark on RATS (Channel A), or asks about evaluating this task. Reports WER (%).
Evaluates a reasoning-based reward model's ability to produce human-aligned preference judgments and optimize visual generation models via reinforcement learning and test-time prompt refinement. Use when the user wants to benchmark on Multimodal Reward Bench 2 (MMRB2), EditReward Bench, GenAI-Bench, ImgEdit-Bench, GEdit-Bench-EN, UniGen (UniGenBench++), PICA-Bench, or asks about evaluating this task. Reports pairwise comparison accuracy.
Measures the proportion of times an LLM selects a stereotypical option over anti-stereotype or unrelated alternatives when prompted implicitly or explicitly. Probes the model's susceptibility to implicit bias and its explicit recognition of stereotypes across demographic categories. Use when the user wants to benchmark on StereoSet, CrowSPairs, or asks about evaluating this task. Reports ratio of stereotypical responses.
Evaluates deep learning models for real-time instance segmentation and ripeness classification of raspberries and punnets in industrial conveyor settings. It probes the model's ability to distinguish between five ripeness grades (OK, Dark, Light, Second, Waste) and background objects under conditions of color similarity and occlusion. Use when the user wants to benchmark on RaspGrade, or asks about evaluating this task. Reports mAP50.
This evaluation probes the transferability and generalization of medical image foundation models pre-trained exclusively on randomized synthetic data. It measures performance across diverse anatomical regions, imaging modalities (CT, MR, X-ray, ultrasound, fundus), and downstream tasks including segmentation, classification, and detection. Use when the user wants to benchmark on TotalSegmentator, CHAOS, LUNA16, INbreast, STARE, DDTI, or asks about evaluating this task. Reports Dice score, AUC.
Evaluates whether dense retrievers and re-rankers can semantically encode and retrieve correct answers to reasoning problems across diverse tasks. It probes the retriever-LLM behavioral gap by testing performance with and without task instructions, and compares full-dataset retrieval against multiple-choice retrieval settings. Use when the user wants to benchmark on RAR-b, or asks about evaluating this task. Reports nDCG@10.
Probes the re-ranking capability of distilled cross-encoder models on passage retrieval tasks. It evaluates ranking quality on in-domain benchmarks (TREC Deep Learning tracks) and out-of-domain generalization across diverse corpora (TIREx framework), while also measuring computational efficiency. Use when the user wants to benchmark on Rank-DistiLLM, TREC DL 2019, TREC DL 2020, TIREx, or asks about evaluating this task. Reports nDCG@10.
Evaluates continual learning methods in online, exemplar-free, and low-exemplar regimes by measuring how well a model retains knowledge of previously seen classes after processing a single pass of sequential data. It specifically tests whether fixed random representations can match or exceed learned representations in these constrained settings. Use when the user wants to benchmark on MNIST, CIFAR10, CIFAR100, TinyImageNet200, miniImageNet100, or asks about evaluating this task. Reports avera...
Evaluates medical imaging models on hand radiographs for rheumatoid arthritis. It probes anatomical structure modeling through bone segmentation, fine-grained lesion detection via bone erosion segmentation, and clinical reasoning through ordinal scoring of erosion and joint space narrowing severity. Use when the user wants to benchmark on RAM-H1200, or asks about evaluating this task. Reports DSC, QWK.
Evaluates the training throughput and network communication efficiency of distributed CNN training frameworks under varying GPU counts and dataset complexities. Probes how well a system mitigates parameter server bottlenecks and scales across multiple concurrent workloads. Use when the user wants to benchmark on ImageNet-1K, ImageNet-22K, or asks about evaluating this task. Reports throughput (images/sec).
Evaluates the perceptual quality of synthesized rakugo speech by comparing it to professional human performances across multiple dimensions, including naturalness, character distinguishability, content understandability, entertainment value, and overall skill level. The benchmark probes whether TTS systems can capture the nuanced performance modeling required for traditional Japanese verbal entertainment. Use when the user wants to benchmark on Misomame, or asks about evaluating this task. Re...
Evaluates deep learning models for spatial precipitation downscaling by measuring both static reconstruction accuracy and dynamic temporal evolution of rainfall patterns. It probes whether models can capture realistic meteorological properties like heavy rain coverage, cluster movement, and transition speeds. Use when the user wants to benchmark on RainNet, or asks about evaluating this task. Reports PEM.
Evaluates deep learning models' ability to forecast global precipitation at multiple lead times (1, 3, 5 days) and estimate same-timestep precipitation using multi-modal satellite and reanalysis data. It probes the model's capacity to handle extreme weather events, class imbalance, and spatial-temporal dependencies in meteorological forecasting. Use when the user wants to benchmark on RainBench, or asks about evaluating this task. Reports Latitude-weighted RMSE.
Evaluates the adversarial robustness and transferability of AI-generated image detectors against crafted perturbations. It probes whether detectors can maintain classification accuracy when faced with white-box and black-box evasion attacks across different perturbation budgets. Use when the user wants to benchmark on RAID, or asks about evaluating this task. Reports F1-score.
Evaluates the factual accuracy and atomic fact alignment of LLMs and RAG systems when answering questions about protein-protein interactions (PPIs) in drug discovery. It probes whether models can correctly identify biological, functional, or physical effects between proteins without hallucinating domain-specific details. Use when the user wants to benchmark on RAGPPI, or asks about evaluating this task. Reports F1 (Cosine similarity of atomic facts).
Evaluates LLM agents' multi-turn decision-making and reasoning capabilities across symbolic planning, risk-sensitive reasoning, and realistic web interaction environments. It probes the agent's ability to complete interactive tasks under noisy or probabilistic feedback while maintaining exploration and training stability. Use when the user wants to benchmark on Bandit, Sokoban, Frozen Lake, WebShop, or asks about evaluating this task. Reports success rate.
Evaluates how Retrieval-Augmented Generation (RAG) systems maintain factual accuracy when exposed to adversarial, harmful, or misleading medical evidence. It probes the model's susceptibility to contextual manipulation and its ability to resist misinformation propagation under varying query framings. Use when the user wants to benchmark on TREC Health Misinformation 2020, TREC Health Misinformation 2021, or asks about evaluating this task. Reports ground-truth alignment rate.
Evaluates retrieval-augmented reasoning systems on their ability to iteratively refine answers using a critique language model. It probes robustness to noisy retrieval, out-of-distribution generalization, and the effectiveness of contrastive critique synthesis over standard self-refinement baselines. Use when the user wants to benchmark on PopQA, TriviaQA, NaturalQuestions, 2WikiMultihopQA, ASQA, HotpotQA, SQuAD, or asks about evaluating this task. Reports accuracy.
This evaluation probes the relationship between retrieval effectiveness and downstream information coverage in RAG systems. It measures how well retrieval models capture required information nuggets and how accurately generated responses cover these nuggets with proper citations. Use when the user wants to benchmark on NeuCLIR24, RAG24, WikiVideo, or asks about evaluating this task. Reports Nugget Coverage.
Evaluates dense optical flow estimation accuracy and generalization across synthetic and real-world driving scenes. It measures pixel-wise displacement error and outlier rates on clean and final passes of benchmark datasets. Use when the user wants to benchmark on Sintel, KITTI, or asks about evaluating this task. Reports EPE.