Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,845
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,321–7,344 of 20,845 skills

Ft Ncfm EvalA

Evaluates the performance and data efficiency of a Vision-Language-Action (VLA) model trained on a synthetically distilled coreset compared to models trained on full datasets. It probes long-horizon manipulation, multi-task skill acquisition, and generalization across spatial, object, goal, and temporal dimensions. Use when the user wants to benchmark on CALVIN, Meta-World, LIBERO, or asks about evaluating this task. Reports Success Rate (SR %), Average Task Completion Length (Avg. Len).

researchpythongo
0
3
Fsl Episodic EvalA

This benchmark evaluates few-shot generalization capability by measuring classification accuracy in an episodic setting where models must recognize novel classes using only a few labeled support examples. It probes the model's ability to adapt quickly to new categories and its robustness to test-time data augmentation. Use when the user wants to benchmark on miniImagenet, tieredImagenet, CUB, Animals, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Fscil EvalA

Probes a model's ability to learn new classes incrementally in a few-shot setting while retaining knowledge of previously learned classes, measuring resistance to catastrophic forgetting. Use when the user wants to benchmark on miniImageNet, CIFAR-100, CUB-200, or asks about evaluating this task. Reports Average accuracy.

researchpythontesting
0
3
Fs Mol EvalA

Few-shot molecular property prediction and regression on a large-scale benchmark with thousands of tasks. Probes generalization across diverse protein targets and varying support set sizes. Use when the user wants to benchmark on FS-Mol, or asks about evaluating this task. Reports ΔAUPRC.

researchpythonperformance
0
3
Frontiermath EvalA

Evaluates advanced mathematical reasoning and experimental problem-solving. It tests whether models can iteratively write and execute Python code to verify hypotheses, refine strategies, and derive correct solutions to expert-level, unsolved math problems. Use when the user wants to benchmark on FrontierMath, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Frogent Drug Design EvalA

Evaluates an end-to-end agentic framework for small-molecule drug design across eight benchmarks spanning the full discovery pipeline, from target identification and knowledge retrieval to virtual screening, interaction profiling, de novo design, and retrosynthetic planning. Use when the user wants to benchmark on Humanity’s Last Exam (HLE), UniProt, Open Targets Platform, ADMETLab 3.0, DAVIS, PLIP, CrossDocked, USPTO-50k, PaRoute, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Frill Noss Esc50 EvalA

Evaluates the quality of lightweight, non-semantic speech embeddings by training simple downstream classifiers on averaged embedding features to perform audio classification tasks. It also measures inference latency on a mobile device to assess real-time suitability for on-device deployment. Use when the user wants to benchmark on NOSS benchmark, ESC-50 (human sounds subset), Mask speech dataset, or asks about evaluating this task. Reports test accuracy.

researchpythontesting
0
3
Freshwiki Article EvalA

Evaluates the ability of LLMs to generate comprehensive, well-organized, and verifiable Wikipedia-like articles from a given topic. It probes outline planning, factual coverage, structural coherence, and source grounding. Use when the user wants to benchmark on FreshWiki, or asks about evaluating this task. Reports ROUGE-1.

researchpythongo
0
3
Freshretailnet 50k EvalA

This benchmark evaluates models' ability to recover latent demand during stockout periods in perishable retail. It probes whether algorithms can disentangle true consumption patterns from supply-induced censoring using hourly temporal data and contextual covariates. Success is measured by prediction accuracy, bias mitigation, and the decoupling of recovered demand from stockout ratios. Use when the user wants to benchmark on FreshRetailNet-50K, or asks about evaluating this task. Reports WAPE.

researchpythongo
0
3
Freqrec EvalA

Evaluates a sequential recommendation model's ability to predict the next item in a user's interaction history by jointly modeling intra-session and inter-session behavioral dynamics. It probes the model's recommendation accuracy, robustness to noisy cross-domain data, and stability under sparse interaction conditions. Use when the user wants to benchmark on Amazon Beauty, Sports & Outdoors, Toys & Games, or asks about evaluating this task. Reports HR@K, NDCG@K.

researchpythongo
0
3
Frenchbench EvalA

Evaluates bilingual French-English language understanding, cultural knowledge, and generation capabilities of LLMs across classification and open-ended tasks. Probes the model's ability to perform few-shot reasoning, factual recall, and text generation in both languages. Use when the user wants to benchmark on FrenchBench, English Benchmarks, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Free Geometry EvalA

Evaluates test-time self-supervised adaptation for feed-forward 3D reconstruction models. It probes the model's ability to refine camera pose estimation and 3D geometry reconstruction on unseen scenes by enforcing cross-view feature consistency without ground-truth labels. Use when the user wants to benchmark on ETH3D, ScanNet++, 7-Scenes, HiROOM, or asks about evaluating this task. Reports AUC@3, F1-score.

researchpythonperformance
0
3
Fredo EvalA

Evaluates few-shot document-level relation extraction by testing a model's ability to identify relations between entity pairs across documents using limited support examples. It specifically probes domain adaptation capabilities, handling of class imbalance, and robustness to NOTA (none-of-the-above) distributions in realistic document-level settings. Use when the user wants to benchmark on FREDo, or asks about evaluating this task. Reports macro F1.

researchpythongo
0
3
Frechet Inception DistanceA

Measures the distance between the feature distributions of real and generated images using a pretrained Inception network. It evaluates both the fidelity and diversity of generated samples by comparing their mean and covariance in the feature space. Use when the user has predictions and gold and needs to compute FID.

researchpythongo
0
3
Frd Optical Fibre TestingA

Evaluates the focal ratio degradation (FRD) of multi-mode optical fibres under automated testing conditions to verify compliance with astronomical instrumentation specifications. It compares automated optical bench measurements against manual ring tests to ensure measurement consistency and accuracy. Use when the user has predictions and gold and needs to compute FRD (Focal Ratio Degradation).

researchpythongo
0
3
Fraudster Group Detection EvalA

Evaluates a model's ability to detect fraudulent reviewer groups by analyzing spatio-temporal co-review patterns. It probes the model's capacity to distinguish genuine groups from coordinated fraudster groups using graph representation learning and temporal modeling. Use when the user wants to benchmark on Yelp, Amazon, or asks about evaluating this task. Reports F1-value.

researchpythongo
0
3
Fraud R1 EvalA

This benchmark evaluates large language models' robustness against multi-round fraud and phishing inducements. It probes whether models can successfully identify and defend against deceptive prompts across five fraud categories under both standard helpful-assistant and role-play settings, while also measuring cross-lingual performance gaps. Use when the user wants to benchmark on Fraud-R1, or asks about evaluating this task. Reports Defense Success Rate (DSR).

researchpythongo
0
3
Fraud Dataset Benchmark EvalA

This benchmark evaluates the robustness of fraud detection models to label noise in training data. It measures how effectively various noise-removal techniques preserve predictive performance when tested on clean, unseen data. The protocol specifically probes a model's ability to mitigate artificially injected label corruption across multiple real-world fraud datasets. Use when the user wants to benchmark on Fraud Dataset Benchmark (FDB), or asks about evaluating this task. Reports ROC-AUC.

researchpythongo
0
3
Framework Throughput EvalA

Evaluates the training throughput and execution efficiency of deep learning frameworks by measuring how quickly they process standard model architectures on a single GPU. It compares PyTorch against TensorFlow, MXNet, CNTK, Chainer, and PaddlePaddle to assess device utilization and runtime optimization. Use when the user has predictions and gold and needs to compute Throughput.

researchpythongo
0
3
Framework Latency EvalA

Evaluates the computational efficiency and hardware utilization of five deep learning frameworks (Caffe, Neon, TensorFlow, Theano, Torch) across standard neural network architectures on CPU and GPU hardware. Use when the user wants to benchmark on MNIST, ImageNet, IMDB, or asks about evaluating this task. Reports forward pass time (ms).

researchpythonperformance
0
3
Frame Auc EvalA

Evaluates a model's ability to detect abnormal human activities by predicting multi-timescale future and past pose trajectories. The framework measures prediction errors across different temporal granularities and combines them to identify anomalous frames. Use when the user wants to benchmark on HR-ShanghaiTech, HR-Avenue, Corridor, or asks about evaluating this task. Reports Frame-AUC.

researchpythonperformance
0
3
Fracture Detection EvalA

Evaluates bone fracture detection and localization in pelvic X-ray images using point-based annotations. It measures image-level classification accuracy and pixel-wise localization precision under clinically relevant false positive rates. Use when the user wants to benchmark on PXR Trauma Registry Dataset, or asks about evaluating this task. Reports AU-ROC, FROC.

researchpythonperformance
0
3
Fractional Follow Up MetricA

Evaluates the sky localization precision and detection sensitivity of gravitational-wave detector networks for multi-messenger follow-up of compact binary mergers. It quantifies how well a network can identify and pinpoint sources within a specific distance and localization area threshold. Use when the user has predictions and gold and needs to compute fractional follow-up metric.

researchpythongo
0
3
Fractaldb Pretrain EvalA

Evaluates the effectiveness of pre-training convolutional neural networks on automatically generated fractal image datasets (FractalDB) compared to natural image pre-training and self-supervised learning, measuring downstream classification accuracy on standard benchmarks. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, ImageNet-100, Places-30, ImageNet-1k, Places-365, Pascal VOC 2012, Omniglot, or asks about evaluating this task. Reports classification accuracy.

researchpythongo
0
3