Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,843
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,561–7,584 of 20,843 skills

Fairness Aware Graph Learning EvalA

This benchmark evaluates the trade-offs between predictive utility and fairness across various graph learning algorithms. It probes how well different methods balance accuracy with demographic parity and equal opportunity constraints on real-world graph-structured data. Use when the user wants to benchmark on 7 real-world datasets, or asks about evaluating this task. Reports Δ_SP.

researchpythongo
0
3
Fairness Aware Gnn EvalA

Evaluates the trade-off between prediction accuracy and statistical fairness for graph neural networks on node classification tasks. It probes how different in-processing and preprocessing methods, backbone architectures, and early stopping conditions affect both standard performance metrics and fairness constraints across synthetic, social, and knowledge graph datasets. Use when the user wants to benchmark on Credit, Bail, Pokec-n, Pokec-z, Pokec-n-Large, Pokec-z-Large, DBpedia, YAGO, Wikida...

researchpythongo
0
3
Fairness Aware Automl EvalA

This evaluation probes an AutoML framework's ability to jointly optimize predictive performance and fairness constraints during pipeline search. It measures how well a multi-criteria genetic algorithm balances accuracy metrics against demographic parity, equalised odds, and ABROCA while simultaneously selecting data and features. Use when the user wants to benchmark on adult, credit-card, portuguese-bank-marketing, or asks about evaluating this task. Reports DP.

researchpythongo
0
3
Fairness And Downstream EvalA

Evaluates demographic fairness and downstream NLU task performance of language models. It measures bias across gender, race, and age using established fairness benchmarks, and verifies that fairness interventions do not degrade accuracy on standard classification and regression tasks. Use when the user wants to benchmark on HolisticBias, WEAT/SEAT, CrowS-Pairs, GLUE, or asks about evaluating this task. Reports Fairscore.

researchpythonperformance
0
3
Fairness Algorithms EvalA

This evaluation probes the trade-off between predictive performance and group fairness across various machine learning pipelines. It systematically compares fairness-unaware baselines against preprocessing and in-training fairness interventions, measuring how well algorithms maintain accuracy while satisfying demographic parity and equalized odds constraints. Use when the user wants to benchmark on Titanic, German, Adult, S-D, S-P, I-D, or asks about evaluating this task. Reports Fair Efficie...

researchpythongo
0
3
Fairmt Bench EvalA

Evaluates the fairness and bias resistance of conversational LLMs in multi-turn dialogue settings. It probes whether models accumulate stereotypes or toxic content across turns, handle implicit bias in context, and maintain safety under various interaction patterns like jailbreaks or misinformation. Use when the user wants to benchmark on FairMT-10K, or asks about evaluating this task. Reports bias ratio.

researchpythongit
0
3
Fairmedqa EvalA

Evaluates large language models for demographic bias in medical question answering by measuring performance disparities across counterfactual clinical vignettes that systematically vary race, sex, and socioeconomic status while preserving clinical outcomes. Use when the user wants to benchmark on FairMedQA, or asks about evaluating this task. Reports accuracy disparity (AD).

researchpythongo
0
3
Fairjob EvalA

Evaluates the fairness and predictive performance of job recommendation models on a real-world tabular dataset. It probes whether models exhibit disparate impact across gender groups when predicting senior job opportunities, while accounting for selection bias and privacy constraints inherent in online advertising systems. Use when the user wants to benchmark on FairJob, or asks about evaluating this task. Reports Demographic Parity (DP).

researchpythongit
0
3
Fairdomain EvalA

Evaluates cross-domain medical image segmentation and classification performance while measuring demographic fairness across gender, race, and ethnicity. It probes whether models maintain equitable accuracy across demographic groups when domain shifts occur between imaging modalities (En face vs. SLO fundus images). Use when the user wants to benchmark on FairDomain-Segmentation, or asks about evaluating this task. Reports ESP (Equity-Scaled Performance).

researchpythongo
0
3
Faircontrast Tabular EvalA

This evaluation probes a model's ability to learn fair representations from tabular data by balancing predictive accuracy with demographic parity. It measures how effectively the model mitigates bias across privileged and unprivileged groups defined by sensitive attributes such as gender or age. Use when the user wants to benchmark on Adult, German Credit, Heritage Health, or asks about evaluating this task. Reports Demographic Parity (DP).

researchpythongo
0
3
Fairco Dynamic Ltr EvalA

This evaluation probes a dynamic learning-to-rank algorithm's ability to balance ranking quality with group-level fairness under position bias. It tests whether the model can maintain high relevance-based ranking performance while actively controlling exposure and impact disparities between predefined item groups over a sequence of user interactions. Use when the user wants to benchmark on Ad Fontes Media Bias (semi-synthetic news), MovieLens-20M, or asks about evaluating this task. Reports a...

researchpythongo
0
3
Fair Weak Supervision EvalA

Evaluates a weak supervision pipeline's ability to mitigate labeling function bias and improve fairness across demographic groups. It measures how well a source bias mitigation method recovers accurate pseudolabels while reducing disparities in prediction rates between privileged and underrepresented groups. Use when the user wants to benchmark on Adult, Bank Marketing, CivilComments, HateXplain, CelebA, UTKFace, WRENCH, or asks about evaluating this task. Reports demographic parity gap ($\De...

researchpythonperformance
0
3
Fair Summm EvalA

Evaluates the fairness of abstractive summarization by measuring distributional alignment across social attributes (e.g., sentiment, gender, party). It quantifies how well generated summaries preserve the proportion of diverse perspectives present in the source text, penalizing underrepresentation of minority viewpoints. Use when the user wants to benchmark on PERSPECTIVESUMM, or asks about evaluating this task. Reports Binary Unfair Rate (BUR).

researchpythongit
0
3
Fair Inference Causal Mediation EvalA

Evaluates whether a predictive model can satisfy fairness constraints defined by causal mediation analysis (NDE/PSE) while maintaining out-of-sample accuracy. It probes the model's ability to isolate and eliminate discriminatory pathways from sensitive attributes to outcomes without relying on fully specified outcome models. Use when the user wants to benchmark on COMPAS, Adult (UCI), or asks about evaluating this task. Reports NDE (odds ratio).

researchpythonperformance
0
3
Fair Income Distribution EvalA

Assesses the fairness of a country's income distribution by comparing its Gini index and quintile income shares against a theoretical benchmark derived from professional sports salary allocations. The benchmark assumes that sports salaries, determined by performance and transparent rules, represent a procedurally and distributively fair standard. Use when the user wants to benchmark on World Bank Income Data, or asks about evaluating this task. Reports percentage deviation.

researchpythonperformance
0
3
Fair Credit Scoring EvalA

This benchmark evaluates the trade-off between algorithmic fairness and financial profitability in credit scoring models. It probes how well various fairness-aware preprocessing, in-processing, and post-processing techniques maintain predictive accuracy and demographic parity while minimizing economic loss for lenders. Use when the user wants to benchmark on german, bene, taiwan, uk, pakdd, gmsc, homecredit, or asks about evaluating this task. Reports Profitability (Profit per EUR).

researchpythongo
0
3
Faetar Benchmark EvalA

Evaluates automatic speech recognition (ASR) models on a highly under-resourced language (Faetar/Franco-Provençal) characterized by noisy field recordings, lack of standard orthography, and inconsistent phonetic transcriptions. Use when the user wants to benchmark on Faetar Benchmark, or asks about evaluating this task. Reports PER.

researchpythongo
0
3
Factuality ScoreA

Evaluates the quality of synthetically generated natural language reports derived from tabular data. It probes factual grounding, narrative coherence, hallucination, and the precise preservation of numerical and temporal information from the source table. Use when the user has predictions and gold and needs to compute factuality_score.

researchpythongo
0
3
Factual Scene Graph Parsing EvalA

This benchmark evaluates a model's ability to parse natural language captions into structured scene graphs that faithfully represent described visual elements. It probes compositional generalization and output consistency by testing parsers on both standard and length-constrained splits, measuring how well generated graph structures align with human-annotated ground truth. Use when the user wants to benchmark on FACTUAL, or asks about evaluating this task. Reports SPICE.

researchpythongo
0
3
Factir EvalA

Evaluates open-domain retrieval and re-ranking systems on real-world fact-checking claims. It probes the ability to retrieve indirect, multifaceted evidence from unstructured web sources to support or refute complex queries involving health, politics, and economics. Use when the user wants to benchmark on FactIR, or asks about evaluating this task. Reports nDCG@k.

researchpythongo
0
3
Factcheck Kg Validation EvalA

Evaluates large language models' ability to validate factual claims in knowledge graphs by classifying triples as true or false. It probes internal knowledge retrieval, retrieval-augmented generation (RAG) with external search results, and multi-model consensus strategies for fact-checking reliability. Use when the user wants to benchmark on FactBench, YAGO, DBpedia, or asks about evaluating this task. Reports Class-wise F1 Score.

researchpythongo
0
3
Fact Checking Arena EvalA

Evaluates LLMs on multi-hop fact-checking by measuring claim extraction, evidence retrieval, and justification quality. Uses an arena-style pairwise comparison framework with LLM judges to rank models across multiple reasoning dimensions. Use when the user wants to benchmark on HOVER, FEVERIOUS, or asks about evaluating this task. Reports Accuracy (%).

researchpython
0
3
Fact Based Oie EvalA

Evaluates Open Information Extraction systems on their ability to correctly extract complete facts from sentences, moving beyond token-level overlap to fact-level exact matching against exhaustive gold synsets. It measures whether a system can identify all surface realizations of a fact and penalizes extractions that contain correct tokens but express incorrect or incomplete facts. Use when the user wants to benchmark on CaRB, or asks about evaluating this task. Reports Precision, Recall, F1 ...

researchpythongo
0
3
Facies Classification EvalA

This benchmark evaluates machine learning models on 3D seismic facies classification, a task critical for geological interpretation. It probes a model's ability to accurately segment and label distinct geological strata from 3D seismic data using both local patch-based and global section-based contextual information. Use when the user wants to benchmark on F3 Block, or asks about evaluating this task. Reports MCA.

researchpythonperformance
0
3