Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

21,245
skills in category
886
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 9,049–9,072 of 21,245 skills

Chainv EvalA

Evaluates the accuracy and inference efficiency of training-free multimodal reasoning methods across diverse vision-language benchmarks. It probes how well atomic visual hint injection reduces redundant reasoning steps while maintaining or improving task performance on math, logic, science, and general visual understanding tasks. Use when the user wants to benchmark on MathVista mini, MathVision, WeMath, MMMU Pro vis, LogicVista, OlympiadBench, VStar, CVBench, ConBench, ChartVQA, SEED-Bench, ...

researchpythongo
0
3
Chain Of Instructions EvalA

Evaluates LLMs on multi-step compositional instruction following, generalization to hard single-step tasks, and multilingual summarization. It probes the model's ability to chain subtask outputs as inputs for subsequent steps and maintain coherence across multiple instructions. Use when the user wants to benchmark on CoI2-test, CoI3-test, BIG-Bench Hard (BBH), Multilingual Summarization, or asks about evaluating this task. Reports Rouge-L.

researchpythonperformance
0
3
Chaic EvalA

Evaluates AI agents' ability to socially perceive and cooperatively assist physically constrained humans in long-horizon indoor and outdoor tasks. It probes cooperative planning, goal inference from egocentric visual input, and emergency response under physical constraints. Use when the user wants to benchmark on CHAIC, or asks about evaluating this task. Reports Transport Rate (TR).

researchpythongo
0
3
Ch4 Detection Intensity EvalA

Evaluates ensemble machine learning models for binary detection of fugitive methane emissions and continuous prediction of their tracer concentration intensity using meteorological data. Use when the user wants to benchmark on HYSPLIT-generated Methane Emission Dataset, or asks about evaluating this task. Reports accuracy.

researchpythongit
0
3
Cgiqa 6k EvalA

Evaluates the ability of no-reference image quality assessment (IQA) models to predict human-perceived quality scores for in-the-wild computer graphics images. It probes how well models capture both distortion artifacts and aesthetic quality in synthetic visual content compared to natural scenes. Use when the user wants to benchmark on CGIQA-6k, CCT-CGI, NBU-CIQAD, LIVE-YT-Gaming, or asks about evaluating this task. Reports SRCC.

researchpythongit
0
3
Cgce EvalA

Evaluates Chinese generative chat models on general knowledge and financial domain tasks, measuring response quality across multiple human-assessed dimensions. It probes the model's ability to handle diverse prompts in mathematics, reasoning, scenario writing, and financial analysis, while assessing the overall quality of the generated Chinese text. Use when the user wants to benchmark on CGCE, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cg Bench EvalA

Evaluates multimodal large language models on clue-grounded audio-visual counting tasks over long videos. It probes the model's ability to integrate audio and visual cues to locate temporal segments and accurately count events, objects, or attributes within those segments. Use when the user wants to benchmark on CG-Bench, or asks about evaluating this task. Reports counting_accuracy.

researchpythongo
0
3
Cfsl Instance EvalA

Evaluates a model's ability to perform continual few-shot learning and recognize specific object instances under varying class counts and corruption levels. It probes instance-level memorization and robustness to noise and occlusion in a streaming episodic setting. Use when the user wants to benchmark on CFSL synthetic images (SlimageNet64), or asks about evaluating this task. Reports accuracy.

researchpythongit
0
3
Cfsl Benchmark EvalA

Evaluates a model's ability to learn sequentially from small, task-specific data batches (continual few-shot learning) without access to prior tasks, measuring sample efficiency and susceptibility to catastrophic forgetting across sequential 5-way 1-shot classification tasks. Use when the user wants to benchmark on Omniglot, SlimImageNet64, or asks about evaluating this task. Reports accuracy.

researchpythontesting
0
3
Cfr Retrieval EvalA

Evaluates an information retrieval system's ability to navigate complex, hierarchical, and temporally-varying regulatory documents. It probes the model's capacity to resolve dense cross-references and versioning conflicts to provide complete and accurate answers. Use when the user wants to benchmark on Code of Federal Regulations (CFR), or asks about evaluating this task. Reports Accuracy (Correct/Complete Answers).

researchpythongo
0
3
Cfq EvalA

This benchmark evaluates compositional generalization in semantic parsing by measuring how well models translate anonymized natural language questions into executable SPARQL queries. It specifically probes the ability to generalize to unseen combinations of logical rules (compounds) while maintaining familiarity with individual rules (atoms), using Maximum Compound Divergence splits to ensure fair yet challenging evaluation. Use when the user wants to benchmark on CFQ, or asks about evaluatin...

researchpythongo
0
3
Cflue EvalA

Evaluates large language models' proficiency in Chinese financial domain knowledge and their ability to perform standard NLP tasks within the financial sector. It probes both factual recall and reasoning via multiple-choice qualification exams, as well as practical application skills like text classification, machine translation, relation extraction, reading comprehension, and text generation. Use when the user wants to benchmark on CFLUE, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cfkgr EvalA

Evaluates a model's ability to perform counterfactual reasoning on knowledge graphs by determining whether a target triple remains plausible after a hypothetical scenario is introduced, while also assessing knowledge retention for unaffected facts. Use when the user wants to benchmark on CFKGR-CoDEx-S, CFKGR-CoDEx-M, CFKGR-CoDEx-L, CFKGR-CoDEx-M*, or asks about evaluating this task. Reports Overall F1-score.

researchpythongo
0
3
Cfinbench EvalA

Evaluates large language models' domain-specific knowledge and reasoning in the Chinese financial context, covering foundational knowledge, professional certifications, practical tasks, and regulatory compliance. Use when the user wants to benchmark on CFinBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cfever EvalA

Evaluates a model's ability to retrieve supporting or refuting evidence from Chinese Wikipedia documents and sentences, and subsequently verify claims by predicting their factual status (Supports, Refutes, or Not Enough Info) based on the retrieved evidence. Use when the user wants to benchmark on CFEVER, or asks about evaluating this task. Reports FEVER Score.

researchpythongo
0
3
Cfdb EvalA

This benchmark evaluates machine learning models' ability to detect fraudulent customer activity by analyzing aggregated behavioral patterns and transaction features at the customer level. It probes anomaly detection and risk profiling capabilities on synthetic, privacy-compliant financial data with highly imbalanced class distributions. Use when the user wants to benchmark on CFDB (Customer-level Fraud Detection Benchmark), or asks about evaluating this task. Reports F1 Score.

researchpythongit
0
3
Ceval EvalA

Evaluates Chinese foundation models' domain knowledge and reasoning capabilities across 52 academic disciplines and four difficulty levels using multiple-choice questions. It probes the models' ability to follow instructions, perform in-context learning, and generate chain-of-thought reasoning in a Chinese language context. Use when the user wants to benchmark on C-EVAL, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cesped Pose Estimation EvalA

This benchmark evaluates supervised deep learning models for predicting 3D particle orientations (rotation matrices) from 2D Cryo-EM micrographs. It assesses both angular prediction accuracy and the downstream quality of 3D structural reconstructions derived from the predicted poses. Use when the user wants to benchmark on CESPED, or asks about evaluating this task. Reports MAnE.

researchpythonshell
0
3
Cesnet Timeseries24 Cl EvalA

This evaluation probes the stability and performance of continual learning algorithms on multivariate time-series forecasting when the underlying data stream is partitioned into tasks with varying temporal granularities. It measures how sensitive forecasting accuracy, catastrophic forgetting, and backward transfer are to the choice of task boundaries and window lengths. Use when the user wants to benchmark on CESNET-Timeseries24, or asks about evaluating this task. Reports Average MSE.

researchpythongo
0
3
Certified Malware Detection EvalA

Evaluates the robustness of malware classifiers against metamorphic evasion attacks and synthetic feature-space perturbations. It probes whether randomized smoothing and majority voting can maintain detection accuracy and recall when executables are structurally altered or corrupted. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports Recall.

researchpython
0
3
Cendol EvalA

This evaluation protocol assesses the language proficiency, generalization capability, and local cultural commonsense reasoning of instruction-tuned LLMs across Indonesian and nine indigenous languages. It probes zero-shot performance on seen and unseen tasks and languages, as well as nuanced understanding of regional proverbs, figures of speech, and story endings. Use when the user wants to benchmark on COPAL-ID, MABL, IndoStoryCloze, MAPS, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cello EvalA

This benchmark evaluates the causal reasoning capabilities of Large Vision-Language Models (LVLMs) across four levels of the Ladder of Causation: discovery, association, intervention, and counterfactual. It probes whether models can correctly identify causal relationships, handle confounding and collider biases, reason about interventions, and generate counterfactual explanations based on visual scenes. Use when the user wants to benchmark on CELLO, or asks about evaluating this task. Reports...

researchpythongo
0
3
Cell Instance Segmentation EvalA

This evaluation probes a model's ability to perform precise instance segmentation on challenging biomedical microscopy images. It specifically tests handling of overlapping, irregularly shaped cells across varying contrast modalities (brightfield, phase-contrast, fluorescence) and object densities. Use when the user wants to benchmark on LIVECell, EVICAN2, ISBI2014, Revvity-25, or asks about evaluating this task. Reports AP (Average Precision).

researchpythongo
0
3
Cebench EvalA

Evaluates vision-language-action (VLA) models on cross-embodiment robotic manipulation tasks, including single-arm, bimanual, and mobile manipulation. It probes spatial reasoning, visual generalization under domain randomization, and the ability to unify navigation and manipulation in a single policy. Use when the user wants to benchmark on CEBench, or asks about evaluating this task. Reports success_rate.

researchpythongo
0
3