Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,041–5,064 of 23,574 skills
Evaluates the quality, phonetic accuracy, and prosodic fidelity of Russian speech datasets and generative models across synthesis, restoration, and denoising tasks. It probes natural intonation, stress accuracy, and audio clarity using standardized subjective ratings and automatic speech quality metrics. Use when the user wants to benchmark on Balalaika (Proposed), M-AILABS Russian, RUSLAN, Russian LibriSpeech, SOVA RuYoutube, Mozilla Common Voice 21.0, or asks about evaluating this task. Rep...
Evaluates a model's ability to extend an existing semantic hierarchy (RuWordNet) by predicting hypernym relationships for novel Russian words using only contextual corpus information, without relying on explicit word definitions. It probes contextual lexical grounding and unsupervised taxonomy extension capabilities specifically for Slavic languages. Use when the user wants to benchmark on RUSSE'2020 Taxonomy Enrichment Test Set, or asks about evaluating this task. Reports MAP.
Evaluates a model's ability to perform Word Sense Induction (WSI) by clustering contextual usages of target words into sense groups without prior sense labels. It probes the system's capability to handle morphological complexity and free word order in Russian across different sense granularities. Use when the user wants to benchmark on wiki-wiki, bts-rnc, active-dict, or asks about evaluating this task. Reports ARI.
Evaluates large language models' ability to solve complex, logic-heavy Chinese natural language puzzles and perform multi-step reasoning on diverse benchmark tasks. It probes the model's capacity for progressive reasoning, self-verification, and adaptability to structured prompt frameworks without manual tuning. Use when the user wants to benchmark on Ruozhiba, BIG-Bench-Hard, or asks about evaluating this task. Reports accuracy.
Evaluates the inference latency and computational runtime of transformer models and individual MLX operations across different hardware platforms (Apple Silicon vs NVIDIA GPU) and input configurations. Use when the user has predictions and gold and needs to compute runtime.
Evaluates the computational running time, speedup gains, and parallel efficiency of median, standard deviation, and full source-finding algorithms on simulated radio interferometric images. Use when the user has predictions and gold and needs to compute Running time (seconds).
Evaluates a model's ability to determine the veracity of social media rumours and classify the discourse stance of replies within a conversation tree. It probes contextual discourse analysis, stance detection, and truthfulness judgment in noisy, interactive text. Use when the user wants to benchmark on RumourEval (SemEval-2017 Task 8), or asks about evaluating this task. Reports classification accuracy, macroaveraged accuracy.
This evaluation probes a language model's ability to perform rule-based logical reasoning on both in-distribution and out-of-distribution tasks. It measures how well the model can apply explicit and implicit logical rules to derive correct answers under strict exact-match conditions. Use when the user wants to benchmark on BigBench Hard (BBH), BigBench Extra Hard (BBEH), ProverQA, or asks about evaluating this task. Reports pass@1 (hard exact match).
This benchmark evaluates long-context language models' ability to retrieve, trace, aggregate, and answer questions across varying context lengths and task complexities. It probes whether models genuinely attend to injected information or rely on parametric knowledge and context copying as sequence length increases. Use when the user wants to benchmark on RULER, or asks about evaluating this task. Reports exact-match accuracy.
Probes a model's continual learning capability in a partially observable, non-stationary synthetic environment based on the Rule 110 cellular automaton. It measures how well capacity-constrained agents adapt to gradual distribution shifts induced by increasing prediction horizons and evolving task parameters. Use when the user wants to benchmark on Rule 110 Prediction Environment, or asks about evaluating this task. Reports online accuracy.
Evaluates whether data-driven reasoning rubrics improve LLM-based trace correctness classification and serve as effective reward signals for reinforcement learning compared to standard LLM judges and verifiable rewards. Use when the user wants to benchmark on SWE-Bench, NuminaMath, NaturalReasoning, or asks about evaluating this task. Reports Balanced Accuracy.
This benchmark evaluates a model's ability to localize specific objects in remote sensing satellite imagery using natural language queries. It probes the model's robustness to scale variations, cluttered backgrounds, and multi-granularity textual descriptions common in aerial/satellite scenes. Use when the user wants to benchmark on RSVGD, or asks about evaluating this task. Reports Pr@0.5.
Evaluates a model's ability to predict RST-style discourse tree structure and nuclearity relations between elementary discourse units (EDUs). It probes both intra-domain and inter-domain generalization of discourse parsing across different text genres (news, instructions, reviews). Use when the user wants to benchmark on RST-DT, Instr-DT, MEGA-DT, Yelp13-DT, or asks about evaluating this task. Reports Parseval (Par.) / RST-Parseval (R-Par.).
This evaluation probes a semi-supervised temporal intrusion detection system's ability to classify network traffic flows as benign or malicious under label scarcity, adversarial contamination, and temporal distribution shifts across heterogeneous cloud environments. Use when the user wants to benchmark on CIC-IDS2017, CSE-CIC-IDS2018, UNSW-NB15, or asks about evaluating this task. Reports detection accuracy.
Evaluates the ability of monocular depth estimation and stereo matching models to reconstruct fine-grained road surface profiles and disparities from high-resolution images. It probes the models' accuracy in capturing micro-level road textures and handling near-to-far distance variations under diverse dynamic conditions. Use when the user wants to benchmark on RSRD-dense, RSRD-sparse, or asks about evaluating this task. Reports Abs Rel.
This benchmark evaluates large language models' ability to perform fine-grained, region-specific semantic reasoning on remote sensing image pairs. It probes localized change comprehension by asking models to answer binary, multiple-choice, and open-ended questions about specific changes (e.g., new construction, vegetation loss) within satellite imagery. Use when the user wants to benchmark on RSRCC, or asks about evaluating this task. Reports Accuracy (%).
Evaluates the optimization capability of evolutionary algorithms on a highly multimodal, nonseparable benchmark function across low (d=5) and high (d=20) dimensional settings. It measures how effectively the algorithm navigates complex, multi-peaked likelihood surfaces to locate the global optimum within a fixed computational budget. Use when the user wants to benchmark on Rotated Schaffers F7 (RSF7), or asks about evaluating this task. Reports mean_max_function_value.
Evaluates vision-language models' ability to generate detailed, semantically accurate captions describing changes between bi-temporal remote sensing image pairs, particularly in disaster scenarios. It probes spatiotemporal reasoning, fine-grained environmental change detection, and long-text generation quality. Use when the user wants to benchmark on RSCC, or asks about evaluating this task. Reports ST5-SCS.
Evaluates the ability of LLM-based evolutionary algorithms to optimize session-based recommendation prompts across multiple objectives (accuracy, diversity, and fairness) simultaneously. Use when the user wants to benchmark on RSBench, or asks about evaluating this task. Reports HV.
Evaluates the zero-shot image classification and text-to-image retrieval capabilities of Vision-Language Models (VLMs) fine-tuned on remote sensing data. It probes the model's ability to generalize to unseen RS scenes and text queries without task-specific fine-tuning, while also measuring resistance to catastrophic forgetting on general-domain benchmarks. Use when the user wants to benchmark on AID, EuroSAT, fMoW, Million-AID, PatternNet, RESISC, RSI-CB, ImageNet-1K, UCM Captions, RSICD, RSI...
Evaluates a model's ability to discover recurring visual patterns in a single image. It measures detection accuracy at both the individual pattern instance level and the whole pattern level against human annotations. Use when the user wants to benchmark on RP-1K, or asks about evaluating this task. Reports RP Instance Recall.
This evaluation probes an LLM's ability to generate factually correct responses and mitigate hallucinations under adaptive retrieval-augmented generation. It measures factual accuracy via LLM-based scoring on open-ended questions and exact-match accuracy on multi-step reasoning yes/no questions. Use when the user wants to benchmark on TruthfulQA, StrategyQA, or asks about evaluating this task. Reports FactScore.
This benchmark probes a model's ability to comprehend hand-drawn, object-centric visual instructions (arrows, circles, colors) and translate them into precise spatiotemporal action plans for robotic manipulation. It evaluates both high-level task reasoning and low-level execution accuracy in cluttered, unseen environments. Use when the user wants to benchmark on RoVI Book dataset, SIMPLER, or asks about evaluating this task. Reports action success rate.
Evaluates the ability of text-to-image models to accurately render specific objects at requested bounding box locations while maintaining prompt fidelity, attribute correctness, and overall aesthetic quality. Use when the user wants to benchmark on ROVI validation set, or asks about evaluating this task. Reports Gen Inst..