Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

23,574
skills in category
983
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 4,945–4,968 of 23,574 skills

Scientific Embedding EvalA

Evaluates static (Word2Vec, FastText) and transformer-based (SciBERT, RoBERTa) embedding models on scientific text using intrinsic (word/sentence similarity) and extrinsic (NER, document classification) tasks. Probes the impact of domain-specific pretraining and sub-word tokenization on representation quality and downstream performance. Use when the user wants to benchmark on UNMSRS, SemEval, Clinical STS 2018, Clinical STS 2019, Conll 2003, CHEMDNER, SciERC, Reuters 12, BioChem 8, or asks ab...

researchpythongo
0
3
Scienceworld EvalA

Evaluates an agent's ability to perform procedural scientific reasoning and navigation within an interactive text-based environment. It probes whether models can execute multi-step experiments (e.g., building circuits, measuring temperatures) rather than just retrieving static facts. Use when the user wants to benchmark on ScienceWorld, or asks about evaluating this task. Reports average_score.

researchpythongo
0
3
Scienceqa EvalA

Evaluates multimodal reasoning and scientific question answering by requiring models to process questions, images, and context to select correct multiple-choice answers. It also probes chain-of-thought reasoning capabilities by measuring the quality of generated explanations and lectures. Use when the user wants to benchmark on ScienceQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Scienceie Keyphrase Relation EvalA

This benchmark evaluates systems on mention-level keyphrase extraction and semantic relation extraction from scientific publications. It probes the model's ability to identify, classify, and link keyphrases across different scientific domains using exact match criteria. Use when the user wants to benchmark on ScienceIE, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Scielo Parallel Corpus EvalA

Evaluates the quality of sentence alignment and machine translation performance on a trilingual scientific article corpus. It probes cross-lingual translation accuracy and structural alignment precision in a specialized academic domain. Use when the user wants to benchmark on Scielo Parallel Corpus, or asks about evaluating this task. Reports BLEU.

researchpythonexpress
0
3
Scicode EvalA

Probes large language models' ability to perform scientific reasoning, domain-specific knowledge recall, and code synthesis on real-world research problems. It evaluates performance on both decomposed subproblems and full main problems under varying conditions of background knowledge and context carry-over. Use when the user wants to benchmark on SciCode, or asks about evaluating this task. Reports pass@1.

researchpythongo
0
3
Scicm ScievalA

Evaluates cross-modality scientific information extraction by jointly predicting named entities, result entities, and relations from both full-text paragraphs and scientific tables. It probes a model's ability to handle long documents, align entities across modalities, and generalize across different scientific domains. Use when the user wants to benchmark on ScICM, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Sciclaimeval EvalA

Evaluates multimodal models' ability to verify scientific claims by classifying them as Supported or Refuted based on cross-modal evidence (tables or figures). It probes visual reasoning, table parsing, and resistance to dataset biases or superficial shortcuts. Use when the user wants to benchmark on SciClaimEval, or asks about evaluating this task. Reports macro-F1.

researchpythongo
0
3
Scibench EvalA

This benchmark evaluates large language models' ability to solve college-level scientific problems across mathematics, chemistry, and physics. It probes multi-step quantitative reasoning, unit conversion, physical derivations, and the efficacy of prompting strategies and external computational tools. Use when the user wants to benchmark on SciBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Sciagent Scientific Reasoning EvalA

Evaluates a multi-agent LLM system's ability to solve high-level, cross-disciplinary scientific problems at Olympiad and frontier-exam levels. It probes formal reasoning, proof generation, symbolic derivation, and chemical modeling through adaptive routing and self-verification. Use when the user wants to benchmark on IMO 2025, IMC 2025, IPhO 2024, IPhO 2025, CPhO 2025, IChO 2025, HLE, or asks about evaluating this task. Reports Olympiad Scoring.

researchpythonperformance
0
3
Sci Verifybench EvalA

Evaluates a model's ability to perform cross-disciplinary scientific verification by judging the correctness or equivalence of proposed answers to scientific problems. It probes domain-specific logical reasoning, handling of complex mathematical/scientific transformations, and robustness to prompt variations. Use when the user wants to benchmark on SCI-VerifyBench, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Scholawrite EvalA

Evaluates an LLM's ability to predict human scholarly writing intentions from a LaTeX draft and to iteratively edit the draft according to those intentions. It measures lexical diversity, topic consistency, and intention coverage across a 100-iteration self-writing process. Use when the user wants to benchmark on SCHOLAWRITE, or asks about evaluating this task. Reports intention coverage.

researchpython
0
3
Scholarly Qald EvalA

Tests natural language interfaces for querying scholarly knowledge graphs (DBLP, ORKG) and hybrid multi-source QA, evaluating question-to-SPARQL translation and answer generation accuracy. Use when the user wants to benchmark on Scholarly QALD, or asks about evaluating this task. Reports Exact Match.

researchpythongo
0
3
Schema To Json EvalA

Evaluates the ability of language models to extract structured information from heterogeneous tables (text, LaTeX, HTML, CSV, XML) using only a human-authored JSON schema as supervision. It probes schema-driven information extraction, testing attribute prediction accuracy across diverse domains and input formats without domain-specific labeled data. Use when the user wants to benchmark on MlTables, ChemTables, DisCoMat, SWDE, or asks about evaluating this task. Reports Table-F1.

researchpythongo
0
3
Schema Guided Dstc8 EvalA

Evaluates zero-shot dialogue state tracking across single and multi-domain conversations. It measures the model's ability to predict intents, extract slot values, and maintain accurate dialogue states over long contexts without prior exposure to unseen service domains. Use when the user wants to benchmark on Schema-Guided Dialogue (DSTC8 Track 4), or asks about evaluating this task. Reports Joint Goal Accuracy.

researchpythongo
0
3
Schain Medical Reasoning EvalA

Evaluates medical vision-language models on disease classification and structured visual reasoning. It probes the model's ability to localize lesions, generate clinically faithful chain-of-thought rationales, and produce accurate diagnostic classifications grounded in visual evidence. Use when the user wants to benchmark on S-Chain, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Scenicrules EvalA

Evaluates autonomous driving agents on their ability to navigate stochastic traffic scenarios while satisfying a hierarchical set of multi-objective specifications. It probes how well agents balance conflicting goals like collision avoidance, road compliance, passenger comfort, and progress under varying priority constraints. Use when the user wants to benchmark on ScenicRules Benchmark, or asks about evaluating this task. Reports Violation Score (VS).

researchpythongo
0
3
Scenefake EvalA

This benchmark evaluates audio forensics models on their ability to discriminate between genuine recordings and audio manipulated via acoustic scene forgery using speech enhancement technologies. It specifically measures threshold-free equal error rate (EER) to assess how well models generalize to unseen attacks without relying on a fixed decision boundary. Use when the user wants to benchmark on SceneFake, or asks about evaluating this task. Reports EER.

researchpythongo
0
3
Scene Smith EvalA

Evaluates text-to-3D indoor scene generation systems on their ability to produce dense, physically plausible, and prompt-faithful environments. It probes both visual realism and simulation-readiness, measuring collision-free layouts and stable physics properties required for robotics policy testing. Use when the user wants to benchmark on SceneSmith Prompt Corpus, or asks about evaluating this task. Reports Realism Win%.

researchpythontesting
0
3
Scene Graph Modification EvalA

Evaluates a model's ability to modify a source scene graph into a target scene graph conditioned on a natural language query. It probes incremental structure expansion, joint node-edge prediction, and the preservation of unmodified graph components during editing. Use when the user wants to benchmark on User Generated, MSCOCO, GCC, RSICD, or asks about evaluating this task. Reports Graph-level accuracy.

researchpythongo
0
3
Scene Graph Generation EvalA

Evaluates scene graph generation models on predicting subject-predicate-object triplets while mitigating long-tailed training biases. It probes zero-shot generalization and graph-level semantic coherence through sentence-to-graph retrieval. Use when the user wants to benchmark on Visual Genome (VG), MS-COCO Caption (VG Overlap), or asks about evaluating this task. Reports mR@K.

researchpythongo
0
3
Scene Bench EvalA

Evaluates the factual consistency and scene graph adherence of text-to-image generation models. It probes whether generated images accurately preserve specified objects and their spatial/relational configurations as defined by input scene graphs, rather than just measuring aesthetic quality or text-image alignment. Use when the user wants to benchmark on Visual Genome (VG) test set, MegaSG, or asks about evaluating this task. Reports SGScore.

researchpython
0
3
Scendi ScoreA

Evaluates the intrinsic diversity of text-to-image generative models by isolating model-driven variation from prompt-driven variation. It uses CLIP embeddings to construct a joint image-text kernel covariance matrix and applies Schur complement decomposition to remove text influence before computing spectral entropy. Use when the user has predictions and gold and needs to compute Scendi score.

researchpythongo
0
3
Scenario Bias Financial Misinfo EvalA

Evaluates how scenario-induced contextual factors (personality, region, identity) and multilingual settings alter LLM judgments on financial misinformation claims, quantifying behavioral bias as the performance shift relative to a neutral baseline. Use when the user wants to benchmark on Multilingual Financial Misinformation Dataset, or asks about evaluating this task. Reports Bias_scen.

researchpythongo
0
3