Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,601–6,624 of 20,836 skills
Probes large language models' ability to answer multiple-choice questions derived from South Korean healthcare professional licensing exams. It evaluates domain-specific medical knowledge, regional clinical guideline adherence, and reasoning capabilities in Korean. Use when the user wants to benchmark on KorMedMCQA, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of cross-lingual embedding models to capture nuanced financial semantics and terminology in low-resource Korean text, specifically measuring how well they align with human judgments of sentence similarity in specialized financial contexts. Use when the user wants to benchmark on KorFinSTS, or asks about evaluating this task. Reports Spearman’s ρ.
Evaluates vision-language models on Korean multimodal comprehension, document/table/chart understanding, and open-ended generation capabilities using translated and newly curated benchmarks. Use when the user wants to benchmark on K-MMBench, K-SEED, K-MMStar, K-DTCBench, K-LLaVA-W, or asks about evaluating this task. Reports accuracy.
Evaluates a training-free interpretability method (Δ-IoU) for detecting false negatives in binary industrial defect detection models. It probes whether post-hoc heatmap intersections can reliably flag 'in-distribution yet confidently wrong' predictions on surface defect datasets. Use when the user wants to benchmark on Kolektor SDD, Kolektor SDD2, or asks about evaluating this task. Reports Recall.
This benchmark evaluates large language models' ability to reason through Japanese national healthcare licensing examinations across ten medical professions. It probes domain-specific clinical knowledge, multimodal image interpretation, and high-stakes decision-making under strict, profession-specific passing criteria. Use when the user wants to benchmark on KokushiMD-10, or asks about evaluating this task. Reports accuracy.
Evaluates machine translation quality for Kokborok (a low-resource Tibeto-Burman language) in both English-to-Kokborok and Kokborok-to-English directions. It probes translation adequacy, fluency, and semantic similarity using both automatic metrics and human ratings. Use when the user wants to benchmark on SMOL Test Set, WMT Test Set (Bible domain), or asks about evaluating this task. Reports BLEU.
This benchmark evaluates the ability of large vision-language models to generate accurate, free-form Korean responses to image-based questions. It probes fine-grained capabilities across perception, reasoning, and safety/bias, specifically testing Korean cultural recognition, OCR, document/table/chart understanding, and hallucination robustness. Use when the user wants to benchmark on KOFFVQA, or asks about evaluating this task. Reports KOFFVQA Score.
Evaluates conversational understanding and response generation capabilities of language models in Korean. It probes dialogue comprehension (classifying topics, emotions, relations, dialog acts, and facts) and response selection (choosing or generating appropriate next utterances across various Korean dialogue contexts). Use when the user wants to benchmark on KoDialogBench, or asks about evaluating this task. Reports accuracy.
Evaluates the accuracy and inherent social bias of LLMs on a culturally adapted Korean multiple-choice question answering benchmark. It probes whether models rely on explicit contextual information versus ingrained cultural stereotypes when answering questions about various social groups. Use when the user wants to benchmark on KoBBQ, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates multistep soft reasoning capabilities of LLMs in long narratives, specifically testing logical deduction, object placement tracking, and team allocation across English and Korean languages. It probes cross-lingual reasoning transfer and the impact of in-context learning strategies like Chain-of-Thought prompting and task-specific hints. Use when the user wants to benchmark on Ko-MuSR, MuSR, or asks about evaluating this task. Reports accuracy.
Evaluates large language models' ability to systematically cover bounded knowledge universes and perform compositional set-based reasoning. It probes three failure stages: completeness (missing knowledge), awareness (failure to identify requirements), and application (incorrect execution) across multiple domains and languages. Use when the user wants to benchmark on KnowledgeBerg, or asks about evaluating this task. Reports Universe F1.
Probes a multimodal LLM's ability to answer visual questions that require external knowledge retrieval. It specifically tests the model's capacity to dynamically decide when to retrieve information and assess the relevance of retrieved documents using self-reflective tokens, without degrading performance on standard visual-only queries. Use when the user wants to benchmark on Encyclopedic-VQA, InfoSeek, or asks about evaluating this task. Reports BERT matching score (BEM), VQA accuracy.
Evaluates an agentic search agent's ability to gather external knowledge and reference images to enhance text-to-image generation for knowledge-intensive, real-world prompts. It measures how well the agent's search-grounded prompts improve visual correctness, text accuracy, faithfulness, and aesthetics compared to direct generation. Use when the user wants to benchmark on KnowGen, or asks about evaluating this task. Reports K-Score.
Evaluates a model's ability to bind natural language knowledge to visual concepts for high-fidelity image reconstruction and customized generation without full retraining. It probes cross-modal knowledge transfer, concept fidelity, and prompt alignment in diffusion-based image generation. Use when the user wants to benchmark on KnowCusBench, or asks about evaluating this task. Reports CLIP-I-Seg.
Evaluates whether large language models rely on memorization versus genuine logical reasoning by measuring performance drops on logically equivalent but locally perturbed Knights and Knaves puzzles. It probes the model's ability to maintain consistent logical deductions when superficial or structural elements of the problem are altered. Use when the user wants to benchmark on Knights and Knaves (K&K), or asks about evaluating this task. Reports LiMem.
Evaluates the accuracy and robustness of an automated extraction pipeline that converts raw Hebrew parliamentary documents into structured metadata, speaker lists, and text. It probes the system's ability to correctly parse dates, identify speakers, extract sentences, and match speaker names to an official database of Knesset members. Use when the user wants to benchmark on Knesset Corpus, or asks about evaluating this task. Reports success_rate.
Evaluates multimodal understanding in Korean language and context across nine academic disciplines. Probes localized knowledge recall, discipline-specific conventions, and the ability to map visual and textual cues to correct answers in Korean institutional settings. Use when the user wants to benchmark on KMMMU, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates large language models' ability to understand and answer expert-level multiple-choice questions in Korean across diverse academic domains. It specifically probes cultural and linguistic alignment, testing whether models can handle native-language nuances and localized knowledge without relying on translated or English-centric training data. Use when the user wants to benchmark on KMMLU, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to classify Korean text into predefined topic categories, testing core semantic understanding and categorization capabilities in Korean. Use when the user wants to benchmark on KLUE-TC, or asks about evaluating this task. Reports Accuracy.
Tests a model's ability to extract relational triples between entities in Korean text, probing structured information extraction capabilities. Use when the user wants to benchmark on KLUE-RE, or asks about evaluating this task. Reports F1.
Evaluates a model's ability to identify and classify named entities (e.g., person, location, organization) within Korean text, testing token-level understanding. Use when the user wants to benchmark on KLUE-NER, or asks about evaluating this task. Reports F1.
Evaluates a model's ability to track and update dialogue state across turns in Korean conversations, testing multi-turn reasoning and slot filling. Use when the user wants to benchmark on KLUE-DST, or asks about evaluating this task. Reports Joint Accuracy.
Evaluates a model's ability to predict syntactic dependency relations between words in Korean sentences, testing grammatical structure understanding. Use when the user wants to benchmark on KLUE-DP, or asks about evaluating this task. Reports LAS.
Evaluates the accuracy and computational efficiency of PCA-based PSF subtraction algorithms for high-contrast astronomical imaging. Specifically, it measures how well the algorithm mitigates speckle noise while recovering planetary signals compared to a reference implementation. Use when the user wants to benchmark on Beta Pictoris ($\beta$ Pic), HR8799, or asks about evaluating this task. Reports SNR.