Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,845
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,297–7,320 of 20,845 skills

Gamefactory EvalA

This evaluation protocol assesses a video generation model's ability to follow discrete and continuous action inputs while maintaining semantic alignment with text prompts and preserving the original model's visual domain. It measures action-following accuracy, camera pose consistency, text-video semantic relevance, and overall video generation quality across in-domain and open-domain scenes. Use when the user wants to benchmark on GF-Minecraft, VPT (Find Cave), or asks about evaluating this ...

researchpython
0
3
Game Of 24 EvalA

Evaluates an LLM's ability to perform algorithmic search and recursive reasoning within a single generation window, without external tree search or iterative prompting. It probes systematic exploration, pruning, and backtracking capabilities in a mathematical constraint satisfaction task. Use when the user wants to benchmark on Game of 24, or asks about evaluating this task. Reports Success rate.

researchpythongo
0
3
Game EvalA

Evaluates a vision-action model's ability to predict actions and scale ratios from video game footage, testing in-distribution and out-of-distribution generalization across 2D and 3D games. Use when the user wants to benchmark on Video Games (In-Distribution & OOD), or asks about evaluating this task. Reports Pearson correlation.

researchpythongo
0
3
Gamayun EvalA

Evaluates multilingual LLM capabilities across general knowledge, reasoning, mathematics, and cultural understanding in English, Russian, and other languages. It probes zero-shot and few-shot performance on standardized benchmarks and custom cultural knowledge tests. Use when the user wants to benchmark on MMLU, GSM8K, MERA, RuBIN, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gaia EvalA

GAIA probes the ability of AI assistants to perform real-world, conceptually simple tasks that require multi-step reasoning, tool use, and multi-modal processing. It measures robustness in practical everyday reasoning and factual validation on questions explicitly designed to be outside the model's training data. Use when the user wants to benchmark on GAIA, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Gai Nerf EvalA

Evaluates the accuracy and generalization of wireless channel prediction models across diverse indoor environments and frequency bands. It probes the model's ability to predict received signal strength (RSSI) and channel state information (CSI) given spatial coordinates and environmental geometry, while testing robustness to physical scene changes and cross-frequency translation. Use when the user wants to benchmark on Our Own Datasets, Argos channel dataset, NewRF simulated data, or asks abo...

researchpythongo
0
3
Gaeleval EvalA

Evaluates LLMs' morphosyntactic competence, machine translation quality, and culturally grounded question-answering abilities in Scottish Gaelic. It probes how well models handle minority language grammar, idiomatic usage, and domain-specific cultural knowledge without relying on English-centric prompting. Use when the user wants to benchmark on GaelEval, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Gadbench EvalA

Evaluates supervised graph anomaly detection capabilities on static attributed graphs. It benchmarks models across transductive and inductive settings, homogeneous and heterogeneous graph structures, and compares traditional tree ensembles with neighbor aggregation against standard and specialized GNNs. Use when the user wants to benchmark on Reddit, Weibo, Amazon, Yelp, T-Fin, Ellip, Tolo, Quest, DGraph, T-Social, or asks about evaluating this task. Reports AUPRC.

researchpythongo
0
3
G4satbench EvalA

This benchmark evaluates the capability of Graph Neural Networks to solve Boolean satisfiability (SAT) problems. It probes whether GNNs can accurately predict formula satisfiability, generate satisfying variable assignments, and identify unsatisfiable cores, while assessing their ability to learn search heuristics from graph-structured logical representations. Use when the user wants to benchmark on G4SATBench, or asks about evaluating this task. Reports classification accuracy.

researchpythongo
0
3
Fysics EvalA

Evaluates multimodal large language models' ability to perceive, reason about, and generate physical attributes and laws from images, videos, and audio. It probes causal physical reasoning, material property mapping, and cross-modal consistency rather than superficial pattern matching. Use when the user wants to benchmark on FysicsEval, or asks about evaluating this task. Reports average score.

researchpythongo
0
3
Fwi EvalA

Evaluates the ability of deep learning models to perform full waveform inversion (FWI) by predicting subsurface velocity maps from seismic data under varying source frequencies and locations. It probes generalization across different source configurations and robustness to noise and missing traces. Use when the user wants to benchmark on FWI-F, FWI-L, FWI-FL, or asks about evaluating this task. Reports L2 relative error.

researchpythonperformance
0
3
Fuximt Xxzh EvalA

Evaluates multilingual machine translation capability for Chinese-involved pairs (xx-zh), measuring translation quality across varying levels of parallel data availability. It probes how well models leverage cross-lingual knowledge transfer and handle data scarcity in low-resource settings. Use when the user wants to benchmark on xx-zh translation pairs, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Futurex EvalA

Evaluates LLM agents' ability to forecast real-world future events under uncertainty. It probes reasoning depth, tool-use/search capability, and temporal validity by requiring models to answer dynamic, live-updated questions before event resolution. Use when the user wants to benchmark on FutureX, or asks about evaluating this task. Reports overall_score.

researchpythongo
0
3
Futurepedia EvalA

This benchmark evaluates multilingual Retrieval-Augmented Generation (RAG) systems across three tasks: monolingual knowledge extraction, cross-lingual knowledge transfer, and multilingual knowledge selection. It probes a model's ability to retrieve and generate answers in eight languages, assess cross-lingual transfer capabilities, and measure selection bias when presented with conflicting answers across languages. Use when the user wants to benchmark on Futurepedia, or asks about evaluating ...

researchpythongo
0
3
Fuss Sound Separation EvalA

Evaluates audio source separation models on mixed reverberant or dry recordings containing 1–4 active sources. It measures reconstruction fidelity and source-counting accuracy across varying mixture complexities. Use when the user wants to benchmark on FUSS, or asks about evaluating this task. Reports SI-SNR.

researchpythonperformance
0
3
Fus Multimodal Robot EvalA

Evaluates a robot policy's ability to ground heterogeneous sensor modalities (vision, touch, sound) into language instructions for zero-shot task execution in partially observable environments. It probes multimodal prompting, compositional reasoning, and the necessity of auxiliary contrastive and language grounding losses. Use when the user wants to benchmark on WidowX Multimodal Teleoperation Dataset, or asks about evaluating this task. Reports task success.

researchpythongo
0
3
Funsd Form Understanding EvalA

Evaluates end-to-end form understanding on noisy scanned documents, covering text detection, optical character recognition, word grouping, semantic entity labeling, and entity linking. Use when the user wants to benchmark on FUNSD, or asks about evaluating this task. Reports F1-score.

researchpython
0
3
Function Calling EvalA

Evaluates an LLM's ability to correctly identify, retrieve, and invoke external APIs or functions based on a user query. It probes zero-shot and multi-turn function-calling capabilities, including handling live vs. non-live APIs, detecting irrelevant queries, and mitigating hallucinations. Use when the user wants to benchmark on BFCL-v3, API-Bank, or asks about evaluating this task. Reports AST.

researchpythongo
0
3
Fun Audio Chat EvalA

Evaluates a large audio language model's capabilities across spoken question answering, audio understanding, speech recognition, function calling, and instruction following. It probes the model's ability to process speech inputs, generate text/speech outputs, and adhere to complex voice instructions while maintaining speech quality and safety. Use when the user wants to benchmark on VoiceBench, OpenAudioBench, UltraEval-Audio, MMAU, MMAU-Pro, MMSU, Librispeech, Common Voice, Speech-ACEBench, ...

researchpythongo
0
3
Full Duplex Bench EvalA

Evaluates real-time interactive behaviors in full-duplex spoken dialogue models. It specifically probes turn-taking, pause handling, backchanneling, and interruption management capabilities without relying on human studies. Use when the user wants to benchmark on Full-Duplex-Bench, or asks about evaluating this task. Reports descriptive metrics.

researchpythongo
0
3
Fujiview Svf EvalA

Evaluates multimodal late-fusion models for predicting scenic visibility (clear, cloudy, perfect, obscured) across short- to medium-term forecasting horizons (+0d to +3d). It probes the model's ability to integrate visual webcam features with meteorological forecasts to handle class imbalance and temporal dynamics in environmental perception. Use when the user wants to benchmark on FujiView, or asks about evaluating this task. Reports accuracy (ACC).

researchpythonperformance
0
3
Fuelcast EvalA

Evaluates the ability of tabular and time-series regression models to predict ship fuel consumption using operational, environmental, and temporal features. It probes how well models leverage in-context learning, weather covariates, and sequential patterns across different vessel types. Use when the user wants to benchmark on FuelCast, or asks about evaluating this task. Reports MAE.

researchpython
0
3
Ftc Ensemble EvalA

Evaluates whether post-hoc ensembling strategies (e.g., greedy selection, top-N, model averaging) improve classification accuracy and uncertainty calibration over single fine-tuned language models. It probes the robustness of combining multiple finetuned classifiers across varying training data sizes (10% vs 100%). Use when the user wants to benchmark on DBpedia, News, SetFit, SST-2, Tweet, IMDB, or asks about evaluating this task. Reports classification error.

researchpythongo
0
3
Ft Speech Asr EvalA

Evaluates automatic speech recognition (ASR) systems on spontaneous, formal parliamentary speech in Danish. It tests both in-domain recognition accuracy and cross-domain transferability between the new FT Speech corpus and the established SBRead corpus. Use when the user wants to benchmark on FT Speech, SBRead, or asks about evaluating this task. Reports WER.

researchpython
0
3