Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,836
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,601–6,624 of 20,836 skills

Kormedmcqa EvalA

Probes large language models' ability to answer multiple-choice questions derived from South Korean healthcare professional licensing exams. It evaluates domain-specific medical knowledge, regional clinical guideline adherence, and reasoning capabilities in Korean. Use when the user wants to benchmark on KorMedMCQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Korfinsts EvalA

Evaluates the ability of cross-lingual embedding models to capture nuanced financial semantics and terminology in low-resource Korean text, specifically measuring how well they align with human judgments of sentence similarity in specialized financial contexts. Use when the user wants to benchmark on KorFinSTS, or asks about evaluating this task. Reports Spearman’s ρ.

researchpythongo
0
3
Korean Vlm Benchmarks EvalA

Evaluates vision-language models on Korean multimodal comprehension, document/table/chart understanding, and open-ended generation capabilities using translated and newly curated benchmarks. Use when the user wants to benchmark on K-MMBench, K-SEED, K-MMStar, K-DTCBench, K-LLaVA-W, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Kolektor Sdd EvalA

Evaluates a training-free interpretability method (Δ-IoU) for detecting false negatives in binary industrial defect detection models. It probes whether post-hoc heatmap intersections can reliably flag 'in-distribution yet confidently wrong' predictions on surface defect datasets. Use when the user wants to benchmark on Kolektor SDD, Kolektor SDD2, or asks about evaluating this task. Reports Recall.

researchpythonrust
0
3
Kokushimd 10 EvalA

This benchmark evaluates large language models' ability to reason through Japanese national healthcare licensing examinations across ten medical professions. It probes domain-specific clinical knowledge, multimodal image interpretation, and high-stakes decision-making under strict, profession-specific passing criteria. Use when the user wants to benchmark on KokushiMD-10, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Kokborok Mt EvalA

Evaluates machine translation quality for Kokborok (a low-resource Tibeto-Burman language) in both English-to-Kokborok and Kokborok-to-English directions. It probes translation adequacy, fluency, and semantic similarity using both automatic metrics and human ratings. Use when the user wants to benchmark on SMOL Test Set, WMT Test Set (Bible domain), or asks about evaluating this task. Reports BLEU.

researchpython
0
3
Koffvqa EvalA

This benchmark evaluates the ability of large vision-language models to generate accurate, free-form Korean responses to image-based questions. It probes fine-grained capabilities across perception, reasoning, and safety/bias, specifically testing Korean cultural recognition, OCR, document/table/chart understanding, and hallucination robustness. Use when the user wants to benchmark on KOFFVQA, or asks about evaluating this task. Reports KOFFVQA Score.

researchpythongo
0
3
Kodiaqbench EvalA

Evaluates conversational understanding and response generation capabilities of language models in Korean. It probes dialogue comprehension (classifying topics, emotions, relations, dialog acts, and facts) and response selection (choosing or generating appropriate next utterances across various Korean dialogue contexts). Use when the user wants to benchmark on KoDialogBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Kobbq EvalA

Evaluates the accuracy and inherent social bias of LLMs on a culturally adapted Korean multiple-choice question answering benchmark. It probes whether models rely on explicit contextual information versus ingrained cultural stereotypes when answering questions about various social groups. Use when the user wants to benchmark on KoBBQ, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ko Musr EvalA

This benchmark evaluates multistep soft reasoning capabilities of LLMs in long narratives, specifically testing logical deduction, object placement tracking, and team allocation across English and Korean languages. It probes cross-lingual reasoning transfer and the impact of in-context learning strategies like Chain-of-Thought prompting and task-specific hints. Use when the user wants to benchmark on Ko-MuSR, MuSR, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Knowledgeberg EvalA

Evaluates large language models' ability to systematically cover bounded knowledge universes and perform compositional set-based reasoning. It probes three failure stages: completeness (missing knowledge), awareness (failure to identify requirements), and application (incorrect execution) across multiple domains and languages. Use when the user wants to benchmark on KnowledgeBerg, or asks about evaluating this task. Reports Universe F1.

researchpythongo
0
3
Knowledge Based Vqa EvalA

Probes a multimodal LLM's ability to answer visual questions that require external knowledge retrieval. It specifically tests the model's capacity to dynamically decide when to retrieve information and assess the relevance of retrieved documents using self-reflective tokens, without degrading performance on standard visual-only queries. Use when the user wants to benchmark on Encyclopedic-VQA, InfoSeek, or asks about evaluating this task. Reports BERT matching score (BEM), VQA accuracy.

researchpythongo
0
3
Knowgen EvalA

Evaluates an agentic search agent's ability to gather external knowledge and reference images to enhance text-to-image generation for knowledge-intensive, real-world prompts. It measures how well the agent's search-grounded prompts improve visual correctness, text accuracy, faithfulness, and aesthetics compared to direct generation. Use when the user wants to benchmark on KnowGen, or asks about evaluating this task. Reports K-Score.

researchpythongo
0
3
Knowcusbench EvalA

Evaluates a model's ability to bind natural language knowledge to visual concepts for high-fidelity image reconstruction and customized generation without full retraining. It probes cross-modal knowledge transfer, concept fidelity, and prompt alignment in diffusion-based image generation. Use when the user wants to benchmark on KnowCusBench, or asks about evaluating this task. Reports CLIP-I-Seg.

researchpythongo
0
3
Knight Knave EvalA

Evaluates whether large language models rely on memorization versus genuine logical reasoning by measuring performance drops on logically equivalent but locally perturbed Knights and Knaves puzzles. It probes the model's ability to maintain consistent logical deductions when superficial or structural elements of the problem are altered. Use when the user wants to benchmark on Knights and Knaves (K&K), or asks about evaluating this task. Reports LiMem.

researchpythongo
0
3
Knesset Corpus Extraction EvalA

Evaluates the accuracy and robustness of an automated extraction pipeline that converts raw Hebrew parliamentary documents into structured metadata, speaker lists, and text. It probes the system's ability to correctly parse dates, identify speakers, extract sentences, and match speaker names to an official database of Knesset members. Use when the user wants to benchmark on Knesset Corpus, or asks about evaluating this task. Reports success_rate.

researchpythondatabase
0
3
Kmmmu EvalA

Evaluates multimodal understanding in Korean language and context across nine academic disciplines. Probes localized knowledge recall, discipline-specific conventions, and the ability to map visual and textual cues to correct answers in Korean institutional settings. Use when the user wants to benchmark on KMMMU, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Kmmlu EvalA

This benchmark evaluates large language models' ability to understand and answer expert-level multiple-choice questions in Korean across diverse academic domains. It specifically probes cultural and linguistic alignment, testing whether models can handle native-language nuances and localized knowledge without relying on translated or English-centric training data. Use when the user wants to benchmark on KMMLU, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Klue Tc EvalA

Evaluates a model's ability to classify Korean text into predefined topic categories, testing core semantic understanding and categorization capabilities in Korean. Use when the user wants to benchmark on KLUE-TC, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Klue Re EvalA

Tests a model's ability to extract relational triples between entities in Korean text, probing structured information extraction capabilities. Use when the user wants to benchmark on KLUE-RE, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Klue Ner EvalA

Evaluates a model's ability to identify and classify named entities (e.g., person, location, organization) within Korean text, testing token-level understanding. Use when the user wants to benchmark on KLUE-NER, or asks about evaluating this task. Reports F1.

researchpythongo
0
3
Klue Dst EvalA

Evaluates a model's ability to track and update dialogue state across turns in Korean conversations, testing multi-turn reasoning and slot filling. Use when the user wants to benchmark on KLUE-DST, or asks about evaluating this task. Reports Joint Accuracy.

researchpythongo
0
3
Klue Dp EvalA

Evaluates a model's ability to predict syntactic dependency relations between words in Korean sentences, testing grammatical structure understanding. Use when the user wants to benchmark on KLUE-DP, or asks about evaluating this task. Reports LAS.

researchpythongo
0
3
Klip Postprocessing EvalA

Evaluates the accuracy and computational efficiency of PCA-based PSF subtraction algorithms for high-contrast astronomical imaging. Specifically, it measures how well the algorithm mitigates speckle noise while recovering planetary signals compared to a reference implementation. Use when the user wants to benchmark on Beta Pictoris ($\beta$ Pic), HR8799, or asks about evaluating this task. Reports SNR.

researchpythongo
0
3