Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,845
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,369–7,392 of 20,845 skills

Foodseg103 EvalA

Evaluates fine-grained semantic segmentation and ingredient localization in food images. It probes a model's ability to handle pixel-wise mask prediction under high appearance variability, long-tailed class distributions, and cross-domain generalization to unseen cuisines. Use when the user wants to benchmark on FoodSeg103, or asks about evaluating this task. Reports mIoU.

researchpythontesting
0
3
Followir EvalA

Evaluates whether information retrieval models can follow complex, long-form instructions derived from TREC narratives to determine document relevance. It probes the model's ability to interpret conditional, negated, and composite relevance criteria rather than relying solely on keyword matching. Use when the user wants to benchmark on Robust04, News21, Core17, or asks about evaluating this task. Reports p-MRR.

researchpythonperformance
0
3
Folktexts EvalA

Evaluates the calibration and predictive accuracy of language models when used as risk scorers for tabular prediction tasks. It probes whether models can accurately quantify outcome uncertainty (calibration) while maintaining discriminative power (AUC) on natural-language versions of tabular datasets. Use when the user wants to benchmark on folktexts, or asks about evaluating this task. Reports ECE.

researchpythongit
0
3
Folktables EvalA

Evaluates how fairness interventions affect predictive accuracy and fairness violations across different geographic regions and time periods. It probes the stability of fairness metrics under distribution shift and the efficacy of pre-processing, in-processing, and post-processing interventions on tabular demographic data. Use when the user wants to benchmark on Folktables (ACS PUMS), or asks about evaluating this task. Reports accuracy.

researchpythongit
0
3
Foice Detection EvalA

Evaluates the ability of state-of-the-art audio deepfake detectors to distinguish real speech from face-to-voice (FOICE) synthesized speech, and assesses how fine-tuning on FOICE data affects robustness against unseen synthesis pipelines like SpeechT5. Use when the user wants to benchmark on FOICE, SpeechT5, or asks about evaluating this task. Reports EER.

researchpythonperformance
0
3
Fogmachine EvalA

Evaluates a discrete-event simulation framework that fuses dynamic scene graphs with urban environments to model hierarchical, interconnected spaces under partial observability. It probes the simulator's capacity to reproduce emergent temporal behaviors, the accuracy of state reconstruction when agent views are sparse, and the computational efficiency of the underlying simulation engine. Use when the user wants to benchmark on FOGMACHINE Scenarios (Bruchsal, Wenningstedt, Trier), or asks abou...

researchpythonnode
0
3
Focustrack EvalA

Evaluates visual object tracking performance specifically for anti-UAV scenarios, probing a model's ability to maintain target localization under abrupt camera motion, extreme scale variations, and small target sizes in thermal infrared imagery. Use when the user wants to benchmark on AntiUAV, AntiUAV410, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Foca Malware Classification EvalA

This evaluation probes a model's ability to classify Android malware by fusing audio and visual representations derived from raw APK binaries. It measures supervised classification performance across multiple malware families and benign samples using standard accuracy and macro-F1 metrics. Use when the user wants to benchmark on CICMalDroid-2020, Mal-Net, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Fma Genre Classification EvalA

Evaluates music information retrieval models on genre classification tasks using a large-scale, open music dataset. It probes the model's ability to map audio tracks to hierarchical genre labels (single-label or multi-label) using raw audio or precomputed features. Use when the user wants to benchmark on FMA (Free Music Archive), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Flying Serving EvalA

Evaluates the runtime performance of an LLM serving engine under bursty, heterogeneous, and long-context workloads. It probes the system's ability to dynamically switch between data and tensor parallelism to optimize latency and throughput while maintaining memory efficiency compared to static and alternative dynamic baselines. Use when the user wants to benchmark on ShareGPT, CodeActInstruct, HumanEval, Synthetic Workloads, or asks about evaluating this task. Reports TTFT.

researchpythonperformance
0
3
Fluke Robustness EvalA

Evaluates how well NLP models maintain performance when subjected to minimal, linguistically-grounded perturbations (e.g., syntactic voice changes, negation, style shifts, geographical/temporal biases) across classification and generation tasks. It probes model brittleness to covariate shifts introduced by natural language modifications rather than adversarial noise. Use when the user wants to benchmark on KnowRef, Few-NERD, GSM8K, IFEval, or asks about evaluating this task. Reports Unrobustn...

researchpythongo
0
3
Fluidlab EvalA

Evaluates the ability of reinforcement learning and trajectory optimization algorithms to control complex, multi-phase fluid systems interacting with rigid bodies. It probes sample efficiency, gradient-based optimization stability, and sim-to-real transfer in high-dimensional, non-smooth fluid dynamics. Use when the user wants to benchmark on FluidLab, or asks about evaluating this task. Reports accumulated reward.

researchpythongo
0
3
Fluidgym EvalA

Evaluates reinforcement learning algorithms for active flow control tasks, measuring their ability to stabilize fluid dynamics and reduce drag or enhance heat transfer. It probes algorithmic robustness, sample efficiency, and the capacity to transfer policies across dimensionalities and domain sizes. Use when the user wants to benchmark on FluidGym, or asks about evaluating this task. Reports mean reward per step.

researchpythongo
0
3
Flue EvalA

Probes French sequence classification capabilities across sentiment analysis, paraphrase identification, and natural language inference. Use when the user wants to benchmark on CLS, PAWSX, XNLI, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Flowxpert Mawi EvalA

Evaluates a network intrusion detection model's ability to classify benign versus malicious traffic flows in real-world IoT environments. It specifically probes robustness to severe class imbalance, feature sparsity mitigation via context-aware embeddings, and temporal generalization across different time periods. Use when the user wants to benchmark on MAWI, or asks about evaluating this task. Reports F1-Score.

researchpythongo
0
3
Flowtransformer EvalA

This evaluation protocol assesses the effectiveness of various transformer-based architectures for flow-based network intrusion detection. It systematically tests different input encodings, transformer blocks, and classification heads across three standard NIDS datasets to determine optimal configurations for accuracy, model size, and inference speed. Use when the user wants to benchmark on NSL-KDD, UNSW-NB15, CSE-CIC-IDS2018, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Flow360 EvalA

Evaluates the accuracy of predicted optical flow fields on spherical 360° video frames, measuring both endpoint displacement and angular deviation. It also assesses egocentric activity recognition performance using these flow features to test rotation-invariant representation learning. Use when the user wants to benchmark on FLOW360, EGOK360, or asks about evaluating this task. Reports EPE.

researchpythongo
0
3
Flores200 Mt EvalA

Evaluates multilingual machine translation quality across 60 languages and 234 translation directions. It specifically probes a model's ability to handle high-, medium-, and low-resource languages while mitigating directional degeneration in symmetric multi-way translation. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports COMET-22.

researchpythongo
0
3
Flores101 Mt EvalA

Evaluates the translation quality of Neural Machine Translation (NMT) models trained on filtered pseudo-parallel corpora. It measures how well few-shot Quality Estimation (QE) based corpus filtering improves MT performance across low-resource and mid-resource language pairs compared to baselines and other filtering methods. Use when the user wants to benchmark on FLORES101, or asks about evaluating this task. Reports BLEU.

researchpythonperformance
0
3
Flores 200 EvalA

Evaluates machine translation quality across 200 languages by measuring meaning preservation and fluency. It compares automatic metrics (spBLEU, chrF++) against calibrated human judgments using the XSTS protocol, while also assessing translation safety/toxicity. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports XSTS.

researchpythongit
0
3
Florence 2 EvalA

Evaluates a unified vision foundation model's zero-shot and fine-tuned capabilities across diverse computer vision tasks. It probes the model's ability to perform image captioning, visual question answering, object detection, referring expression comprehension, and semantic segmentation using a single sequence-to-sequence architecture. Use when the user wants to benchmark on COCO, Flickr30k, RefCOCO/+/g, VQAv2, ADE20K, or asks about evaluating this task. Reports CIDEr.

researchpythongo
0
3
FlopsA

This protocol evaluates the computational throughput and real-time efficiency of embedded CPU and GPU platforms. It measures peak floating-point operations per second (FLOPS) using a controlled matrix rotation kernel, and assesses practical system performance via end-to-end inference latency and power consumption on a robotic vision pipeline. Use when the user has predictions and gold and needs to compute FLOPS.

researchpythongo
0
3
Flood Forecasting EvalA

Evaluates a model's ability to predict river water levels and forecast floods using spatiotemporal radar precipitation data. It probes the model's accuracy across multiple forecasting lead times (2h to 12h) and its robustness in capturing extreme hydrological events compared to baseline and deep learning models. Use when the user wants to benchmark on Goslar, Göttingen, or asks about evaluating this task. Reports NSE.

researchpythongo
0
3
Flm Audio EvalA

Evaluates native full-duplex audio-language models on speech understanding, speech generation, and real-time conversational capabilities. It measures how well models handle asynchronous text-audio streams, responsiveness to interruptions, and overall dialogue quality compared to specialized ASR/TTS systems and other full-duplex chatbots. Use when the user wants to benchmark on Fleurs-zh, LibriSpeech-clean, LlamaQuestions, Seed-TTS-en, Seed-TTS-zh, Custom Chinese Speech Instruction-Following S...

researchpythontesting
0
3