Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,843
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,537–7,560 of 20,843 skills

Fastat Benchmark EvalA

This benchmark evaluates the adversarial robustness and computational efficiency of Fast Adversarial Training (FastAT) methods. It measures how well models maintain accuracy under strong adversarial attacks (PGD, AutoAttack, CR Attack) while tracking training time and memory usage, ensuring fair comparison by controlling architecture, training settings, and data sources. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny-ImageNet, or asks about evaluating this task. Reports Aut...

researchpythongo
0
3
Fast Vgan Vc EvalA

Evaluates a GAN-based voice conversion model's ability to transfer speaker timbre while explicitly controlling prosodic features like F0 and duration. It tests static prosodic manipulation, dynamic expressive transfer without expressive training data, and real-time inference efficiency. Use when the user wants to benchmark on VCTK Corpus, Expresso Dataset, or asks about evaluating this task. Reports timbre transfer quality.

researchpythonexpress
0
3
Fashionpedia EvalA

Evaluates joint instance segmentation and fine-grained attribute localization on fashion apparel. It measures how well a model can detect objects, segment them accurately, and correctly assign multiple localized attributes to each instance. Use when the user wants to benchmark on Fashionpedia, or asks about evaluating this task. Reports AP_{IoU + F_1}.

researchpythongo
0
3
Fashion Ner El EvalA

Evaluates a BERT-based Named Entity Recognition pipeline and a binary classifier for candidate entity disambiguation on fashion product descriptions. It probes the model's ability to extract attribute mentions (e.g., material, color) and correctly link them to a knowledge graph ontology under severe data scarcity. Use when the user wants to benchmark on Fashion Product Descriptions (In-house), Fashion EL Disambiguation Dataset, or asks about evaluating this task. Reports f1-score.

researchpythongo
0
3
Fashion Compatibility EvalA

Evaluates a model's ability to score the visual-semantic compatibility of items within a complete outfit, and to recommend a missing item that best completes a partial outfit. It probes the model's capacity to learn non-transitive, type-aware relationships across different fashion categories. Use when the user wants to benchmark on Maryland Polyvore, Polyvore Outfits, Polyvore Outfits-D, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Fara 7b Agentic EvalA

This evaluation probes the agentic capabilities of computer-use models by measuring their ability to complete multi-step web browsing and task-completion tasks on live websites. It assesses both functional success rates and operational efficiency, including token usage, cost, and interaction length. Use when the user wants to benchmark on WebVoyager, Online-Mind2Web, DeepShop, WebTailBench, ScreenSpot, or asks about evaluating this task. Reports success rate.

researchpythongo
0
3
Far3det EvalA

Evaluates 3D object detection performance specifically in the far-field range (50-80m), highlighting the limitations of fixed-threshold metrics and sparse lidar data. It probes how well models detect distant objects using lidar, RGB, or fused modalities under adaptive distance-aware tolerance thresholds. Use when the user wants to benchmark on nuScenes, Far nuScenes, or asks about evaluating this task. Reports 3D mAP.

researchpythonperformance
0
3
Fanstore Io EvalA

Evaluates the I/O throughput and bandwidth of a distributed runtime file system (FanStore) across varying node counts and file sizes, comparing it against local SSDs, FUSE, and shared file systems like Lustre. Use when the user wants to benchmark on ImageNet-1k, SRGAN, FRNN, Custom Synthetic Benchmark, or asks about evaluating this task. Reports bandwidth (MB/s), throughput (files/s).

researchpythonnode
0
3
Fanns Benchmark EvalA

Evaluates the accuracy and efficiency of filtered approximate nearest neighbor search (FANNS) algorithms on high-dimensional transformer-based embeddings. It measures how well different indexing methods maintain recall under various real-world attribute filtering constraints while scaling to millions of vectors. Use when the user wants to benchmark on arxiv-for-fanns-medium, arxiv-for-fanns-large, or asks about evaluating this task. Reports recall@10.

researchpythongo
0
3
Fanar20 Benchmarks EvalA

Evaluates bilingual language understanding, reasoning, and instruction-following capabilities of generative AI models. It probes general knowledge, commonsense reasoning, and culturally aligned Arabic comprehension across multiple standard and custom benchmarks. Use when the user wants to benchmark on English Benchmarks (MMLU, HellaSwag, ARC-Challenge, PIQA, Winogrande), OALL v1, or asks about evaluating this task. Reports English Avg., Arabic Avg..

researchpythongo
0
3
Famma EvalA

Evaluates multimodal large language models on financial domain reasoning, including calculation-heavy arithmetic problems and knowledge-intensive non-arithmetic questions. It probes cross-lingual capabilities and robustness to data contamination by testing on both textbook-derived and expert-crafted live questions. Use when the user wants to benchmark on FAMMA-Basic, FAMMA-LivePro, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Falserject EvalA

Evaluates an LLM's tendency to over-refuse benign prompts that merely appear harmful. It probes the model's ability to distinguish safe from unsafe contexts in controversial queries and provide helpful, context-aware responses instead of unnecessary refusals. Use when the user wants to benchmark on FalseReject, or asks about evaluating this task. Reports over-refusal.

researchpythongo
0
3
Falcon EvalA

Evaluates the efficiency and accuracy of homomorphically encrypted convolution operations and end-to-end private inference networks. It measures communication overhead, inference latency, and classification accuracy under simulated WAN/LAN bandwidths and varying polynomial degrees. Use when the user wants to benchmark on CIFAR-10, CIFAR-100, Tiny ImageNet, or asks about evaluating this task. Reports latency.

researchpythongo
0
3
Fakeclue Loki EvalA

Evaluates large multimodal models on synthetic image detection and artifact explanation. It probes the model's ability to classify images as real or fake and generate natural language explanations for specific visual artifacts. Use when the user wants to benchmark on FakeClue, LOKI, or asks about evaluating this task. Reports Acc.

researchpythongo
0
3
Fakebench EvalA

Evaluates large multimodal models on explainable fake image detection across closed-ended classification and open-ended reasoning tasks. It probes the models' ability to accurately classify image authenticity and generate evidence-based, interpretable justifications using visual and textual forensic cues. Use when the user wants to benchmark on FakeBench, or asks about evaluating this task. Reports Accuracy (ACC).

researchpythongo
0
3
Fake Voice Detection EvalA

Evaluates the robustness and cross-domain generalization of fake voice detectors against 17 state-of-the-art generators (TTS, voice conversion, audio reconstruction) using a one-to-one protocol. It quantifies both generator quality and detector effectiveness through composite scores to expose method-specific vulnerabilities that aggregated benchmarks typically mask. Use when the user wants to benchmark on LibriSpeech (test-clean), ASVspoof-21LA, ASVspoof-21DF, ASVspoof-5, Fake or Real (FoR), ...

researchpythonperformance
0
3
Fake Video Detection EvalA

Evaluates the generalization of existing fake image detection models to synthetic videos generated by diffusion models. Probes whether static image-based forgery detectors can identify AI-generated video content when reduced to single frames. Use when the user wants to benchmark on VidProM, DVSC2023, or asks about evaluating this task. Reports Accuracy.

researchpython
0
3
Faithfulness EvalA

Evaluates the faithfulness and causal alignment of chain-of-thought reasoning in LLMs by measuring how much the final answer depends on the generated reasoning steps versus the original question. It uses causal mediation analysis to compute indirect and direct effects, and a simulator-based metric to quantify rationale faithfulness. Use when the user wants to benchmark on StrategyQA, GSM8K, Causal Understanding, Quarel, OpenBookQA, QASC, or asks about evaluating this task. Reports Controlled ...

researchpythongo
0
3
Fairx Gcig EvalA

Evaluates whether a model maintains predictive utility while achieving procedural fairness (explanation invariance across protected groups) and outcome fairness. It probes the alignment between equalized odds and group-level feature attribution consistency. Use when the user wants to benchmark on Adult, German Credit, COMPAS, Bank Marketing, or asks about evaluating this task. Reports GCIG.

researchpythongo
0
3
FairpfnevalA

Evaluates a model's ability to mitigate the causal and counterfactual effects of protected attributes on predictions while maintaining predictive accuracy, without requiring explicit causal graph knowledge. Use when the user wants to benchmark on Synthetic Causal Case Studies, Law School Admissions, Adult Census Income, or asks about evaluating this task. Reports Total Causal Effect (TCE/TeE).

researchpythonperformance
0
3
Fairness Sequential EvalA

Evaluates sequential decision policies under simulated historical and measurement bias to measure how accounting for unrealized outcomes affects fairness disparities and cumulative utility. It probes whether uncertainty-aware exploration mitigates selection rate differences and false positive rate parity violations without sacrificing profit. Use when the user wants to benchmark on Synthetic Sequential Simulation, or asks about evaluating this task. Reports selection rate difference.

researchpythongo
0
3
Fairness Recourse Subgroup EvalA

Evaluates the fairness of algorithmic recourse across demographic subgroups by ranking them according to various counterfactual-based fairness definitions. It probes whether different recourse fairness metrics capture distinct aspects of bias, actionability constraints, and subgroup granularity in real-world decision systems. Use when the user wants to benchmark on Adult, or asks about evaluating this task. Reports unfairness_score.

researchpythongo
0
3
Fairness Explanation EvalA

This benchmark evaluates the faithfulness and utility of counterfactual explanations in recommendation systems. It measures how effectively generated explanations identify fairness-disparaging attributes by iteratively erasing them and observing the resulting impact on recommendation accuracy and item exposure inequality. Use when the user wants to benchmark on Yelp, Douban Movie, Last-FM, or asks about evaluating this task. Reports NDCG@K.

researchpythonperformance
0
3
Fairness Comparison EvalA

Evaluates the comparative performance of fairness-enhancing machine learning interventions across multiple datasets. It probes how different algorithmic strategies trade off predictive accuracy against a comprehensive set of fairness metrics under standardized preprocessing and fixed train-test splits. Use when the user wants to benchmark on Standardized benchmark datasets (unspecified in excerpt), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3