Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

21,231
skills in category
885
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 8,881–8,904 of 21,231 skills

Clinical Outcome Prediction EvalA

Evaluates a model's ability to predict clinical outcomes (mortality and length of stay) by fusing structured ICU time-series data with medical entities extracted from clinical notes. Use when the user wants to benchmark on Clinical ICU dataset (unspecified), or asks about evaluating this task. Reports AUROC.

researchpython
0
3
Clinical Note Understanding EvalA

Evaluates NLP models on hierarchical clinical reasoning tasks, including SOAP section segmentation, diagnostic inference via assessment-plan relation labeling, and clinical summarization through problem/action plan extraction. Use when the user wants to benchmark on Clinical Progress Notes (MIMIC-III), or asks about evaluating this task. Reports Cohen's Kappa.

researchpythongo
0
3
Clinical Note Scoring EvalA

Evaluates the capability of transformer-based models to automatically assign numerical scores to clinical patient notes. It probes how masked language modeling pretraining and pseudo-labeling strategies improve scoring performance across different model architectures. Use when the user wants to benchmark on Unspecified clinical patient notes dataset, or asks about evaluating this task. Reports CV Score.

researchpythonperformance
0
3
Clinical Note EvalA

This evaluation probes the clinical reasoning, safety, and instruction-following capabilities of large language models on real-world medical datasets. It measures how well models generate accurate and appropriate responses to clinical prompts compared to baseline systems and GPT-3.5-turbo. Use when the user wants to benchmark on MIMIC-III, MIMIC-IV, i2b2, MTSamples, CASI (AE), CASI (CR), DisCQ, or asks about evaluating this task. Reports scores.

researchpythongo
0
3
Clinical Ner EvalA

Evaluates language models' ability to identify and classify standardized medical entities (e.g., diseases, drugs, procedures, genes) in unstructured clinical text. It probes sequence labeling performance under strict terminology standardization (OMOP CDM) to ensure interoperability across diverse healthcare datasets. Use when the user wants to benchmark on NCBI Disease corpus, CHIA, BC5CDR, BIORED, or asks about evaluating this task. Reports Macro Average F1-score (token-based).

researchpythongo
0
3
Clinical Modernbert EvalA

Evaluates a biomedical language model on short- and long-context clinical NLP tasks including classification, named entity recognition, and retrieval. It also measures pre-training masked language modeling accuracy and measures inference efficiency under varying computational loads. Use when the user wants to benchmark on EHR-Prediction (MIMIC-IV ED), MedNER, Pubmed-NCT, PMC-Retrieval, i2b2 2006, i2b2 2010, i2b2 2012, i2b2 2014, or asks about evaluating this task. Reports top-k accuracy.

researchpythongo
0
3
Clinical Field Recovery EvalA

Evaluates sequential question-selection strategies for recovering target clinical fields from synthetic patient responses under a fixed interaction budget. It probes how well adaptive versus fixed questioning policies handle varying patient communication styles to maximize information coverage within conversational constraints. Use when the user wants to benchmark on Clinical Psychiatric Intake Vignette Benchmark, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Clinical Decision MetricsA

Evaluates LLMs on real-world clinical decision-making across three dimensions: effectiveness (accuracy on closed-ended questions), efficiency (proportion of correct, non-redundant reasoning steps), and explainability (quality of clinical reasoning/answers via human or LLM judges). Use when the user has predictions and gold and needs to compute Accuracy.

researchpythongo
0
3
Clinical Assertion Detection EvalA

Evaluates a model's ability to classify the assertion status of medical entities in clinical text across six categories: present, absent, possible, hypothetical, conditional, and associated with someone else. It benchmarks fine-tuned LLMs, transformer classifiers, rule-based systems, and commercial APIs to measure domain-specific clinical NLP performance. Use when the user wants to benchmark on i2b2 2010, or asks about evaluating this task. Reports weighted avg performance.

researchpythongo
0
3
Clinconsensus EvalA

Evaluates Chinese medical LLMs on their ability to generate clinically usable, consistent, and safe responses across diverse specialties and difficulty levels. It probes reasoning depth, evidence integration, and longitudinal follow-up rather than raw factual accuracy. Use when the user wants to benchmark on ClinConsensus, or asks about evaluating this task. Reports CACS@7.

researchpythongit
0
3
Clinb EvalA

Evaluates foundational models on climate intelligence by testing their ability to generate long-form, evidence-grounded answers with accurate citations and relevant multimodal content. It probes knowledge synthesis, hallucination rates in references and images, and alignment with expert-curated quality rubrics. Use when the user wants to benchmark on CLINB, or asks about evaluating this task. Reports ELO score.

researchpythongo
0
3
Climdetect EvalA

This benchmark evaluates machine learning models' ability to detect and attribute human-induced climate change signals from daily spatial climate fields. It specifically probes how well models can predict Annual Global Mean Temperature (AGMT) using surface temperature, humidity, and precipitation data, and assesses their sensitivity in identifying the year when climate change signals robustly emerge from natural variability. Use when the user wants to benchmark on ClimDetect, or asks about ev...

researchpythongo
0
3
Climb EvalA

Evaluates clinical foundation models across diverse medical modalities (imaging, time series, graphs, text) using multitask pretraining, few-shot transfer, and multimodal fusion. It probes model robustness on understudied tasks, adaptation to limited labeled data, and integration of heterogeneous clinical signals for prognosis. Use when the user wants to benchmark on CLIMB, or asks about evaluating this task. Reports balanced AUC.

researchpythongo
0
3
Climax Letkf EvalA

Evaluates the stability and error covariance representation of an AI-based weather prediction model (ClimaX) when integrated into an ensemble data assimilation system (LETKF). It probes the model's ability to generate physically consistent ensemble forecasts, capture flow-dependent error growth, and propagate observation information to unobserved variables without filter divergence. Use when the user wants to benchmark on WeatherBench, or asks about evaluating this task. Reports RMSE.

researchpythongo
0
3
Climateiqa EvalA

This benchmark evaluates vision-language models on meteorological heatmap analysis, probing their ability to perform spatial localization, color semantics understanding, and anomaly detection through visual question answering. It tests four distinct capabilities: verifying statements about anomalies, enumerating affected regions, geo-indexing precise coordinates, and generating descriptive analyses. Use when the user wants to benchmark on ClimateIQA, or asks about evaluating this task. Report...

researchpythongo
0
3
Climategpt EvalA

Evaluates large language models on climate-specific knowledge, reasoning, and fact-verification, alongside general domain benchmarks for commonsense reasoning and world knowledge. It also tests multilingual capability via cascaded machine translation on an Arabic exam dataset. Use when the user wants to benchmark on ClimaBench, Pira 2.0 MCQ, Exeter Misinformation, HellaSwag, PIQA, OpenBookQA, WinoGrande, MMLU, EXAMS (Arabic), or asks about evaluating this task. Reports Acc.

researchpythongo
0
3
Climatecheck EvalA

Tests the ability to retrieve relevant scholarly abstracts for climate change claims from social media and classify the relationship between claims and abstracts as supporting, refuting, or inconclusive. Use when the user wants to benchmark on ClimateCheck, or asks about evaluating this task. Reports Recall@10.

researchpythongo
0
3
Climatecause EvalA

Evaluates large language models' ability to infer correlation directions between event pairs and identify causal chain structures (membership and node position) from climate science text, including implicit and nested causal relations. Use when the user wants to benchmark on ClimateCause, or asks about evaluating this task. Reports F1.

researchpythonnode
0
3
Climatebench M EvalA

Evaluates AI models on multi-modal climate data across three tasks: tensor time-series forecasting, extreme weather anomaly detection, and crop classification from satellite imagery. It probes spatial-temporal alignment, generative data synthesis, and robustness to highly imbalanced rare events. Use when the user wants to benchmark on ClimateBench-M, or asks about evaluating this task. Reports Mean Absolute Error (MAE).

researchpythongo
0
3
Climate Segmentation EvalA

This evaluation probes pixel-level weather pattern segmentation (atmospheric rivers and tropical cyclones) from multi-channel climate data. It measures both segmentation accuracy and exascale training throughput/scaling efficiency across different network architectures and hardware configurations. Use when the user wants to benchmark on Climate weather pattern dataset, or asks about evaluating this task. Reports IoU.

researchpythonnode
0
3
Climate Ood Robustness EvalA

Evaluates the out-of-distribution robustness of climate emulators under temporal extrapolation and cross-scenario forcing shifts. It probes whether models trained on historical climate data can accurately generalize to novel future regimes and extreme emission pathways without seeing them during training. Use when the user wants to benchmark on ClimateSet / CMIP6 GCM outputs, or asks about evaluating this task. Reports LL-RMSE.

researchpythonperformance
0
3
Climate Fever EvalA

This evaluation probes a model's ability to verify scientific claims in a binary classification setting, specifically testing out-of-domain generalization. It measures performance on Supported vs. Refuted labels, emphasizing robustness when applied to climate-related claims outside the training distribution. Use when the user wants to benchmark on CLIMATE-FEVER, or asks about evaluating this task. Reports Balanced Accuracy.

researchpythongo
0
3
Climate Eval EvalA

This benchmark evaluates open-source large language models on their ability to understand, classify, and reason about climate-related discourse. It probes capabilities across text classification, stance detection, claim verification, misinformation detection, and named entity recognition using real-world news, corporate reports, social media, and scientific abstracts. Use when the user wants to benchmark on Guardian Climate News Corpus, Climate-Stance, Climate-FEVER, Climate-Change NER, Net-Z...

researchpythongo
0
3
Climate Downscaling EvalA

Evaluates the ability of deep learning models to perform super-resolution (downscaling) on meteorological surface variables across different spatial resolutions and climate datasets. It probes spatial reconstruction accuracy, structural fidelity, and zero-shot generalization capability in Earth system modeling. Use when the user wants to benchmark on ERA5, BARRA-SY, or asks about evaluating this task. Reports RMSE.

researchpythonaws
0
3