Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 6,673–6,696 of 20,827 skills
Evaluates a model's ability to jointly recognize individual actions, social group activities, and global crowd-level activities in crowded panoramic scenes. It probes multi-granular activity recognition and hierarchical graph-based scene understanding. Use when the user wants to benchmark on JRDB-PAR, or asks about evaluating this task. Reports Overall F1 ($\mathcal{F}_a$).
Evaluates the quality of computed game-theoretic strategies in the word-guessing game Jotto by measuring expected guesses against a benchmark opponent and in self-play, as well as the equilibrium approximation error (epsilon). Use when the user wants to benchmark on Jotto (2-5 letter variants), or asks about evaluating this task. Reports epsilon.
Evaluates a layer-granular DNN offloading framework by measuring inference latency and energy consumption across standard discriminative, generative, and autoencoder neural network architectures on mobile-cloud setups. Use when the user wants to benchmark on JointDNN Deep Architecture Benchmarks (AlexNet, OverFeat, VGG16, Deep Speech, ResNet, NiN, Chair, Pix2Pix), or asks about evaluating this task. Reports latency.
Evaluates models' ability to detect physical face spoofing attacks and digital face forgeries using visual appearance and physiological rPPG cues. It measures cross-domain generalization and compares separate versus joint multi-task learning protocols. Use when the user wants to benchmark on SiW, 3DMAD, HKBU-MarsV2, MSU-MFSD, 3DMask, ROSE-Youtu, FaceForensics++, DFDC, CelebDFv2, or asks about evaluating this task. Reports AUC, EER.
Evaluates the ability of a genetic algorithm to evolve a navigation program for an artificial ant to traverse a complex, toroidal grid trail with gaps and high-difficulty sections. The benchmark measures how well the evolved program generalizes beyond the standard Santa Fe trail into a chaotic extended sector. Use when the user wants to benchmark on John Muir Ant Problem, or asks about evaluating this task. Reports score.
Evaluates a model's ability to normalize job titles by mapping them to standardized ESCO occupation labels using semantic similarity. It probes the model's capacity to handle hierarchical occupational taxonomies and filter out irrelevant contextual tokens like locations. Use when the user wants to benchmark on JobBERT Vacancy Titles, or asks about evaluating this task. Reports MRR.
Evaluates multimodal language models' ability to perform integrated visual-textual reasoning on Japanese-language tasks where questions and reference images are combined into a single composite image. It specifically probes OCR capabilities, visual perception, and cross-modal alignment in a multilingual context. Use when the user wants to benchmark on JMMMU-Pro, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates large multimodal models' ability to understand Japanese-language visual content and answer questions across multiple disciplines. It specifically probes the gap between general language translation capabilities (culture-agnostic subset) and deep cultural knowledge (culture-specific subset), revealing how models handle language variation bias and culturally grounded reasoning. Use when the user wants to benchmark on JMMMU, or asks about evaluating this task. Reports ac...
Evaluates Japanese biomedical large language models across five tasks: multiple-choice question answering, named entity recognition, machine translation, document classification, and semantic text similarity. It probes domain-specific knowledge, multilingual comprehension, and in-context learning capabilities. Use when the user wants to benchmark on JMedBench, or asks about evaluating this task. Reports F1-entity, Accuracy.
Evaluates spatial-temporal reasoning and hardware-agnostic robotic manipulation skills through a structured jigsaw puzzle assembly protocol. It measures vision-based segmentation, object recognition, pick planning success, and motion planning efficiency across three progressively complex physical tasks. Use when the user wants to benchmark on Jigsaw Manipulation Benchmark, or asks about evaluating this task. Reports Task Score.
Evaluates object detection, recognition, and precise tiling assembly capabilities for modular robotic manipulation. It tests the pipeline's ability to segment oddly shaped pieces, recognize their identity, and assemble them with high spatial accuracy. Use when the user wants to benchmark on Jigsaw puzzle set, or asks about evaluating this task. Reports score.
Evaluates the trade-off between classification performance (AUC) and model resilience (robustness to Monte Carlo simulation variations) in quark/gluon and top-quark jet tagging. It probes whether complex neural architectures generalize better to different physics simulators compared to simpler, physics-informed models. Use when the user wants to benchmark on Pythia 8 / Herwig 7 Jet Samples, or asks about evaluating this task. Reports AUC.
Evaluates ultra-low-latency supervised classification of particle physics jet signatures on edge hardware. It probes the ability to distinguish rare boson/top-quark jets from common quark/gluon jets under strict microsecond latency and pipeline interval constraints. Use when the user wants to benchmark on LHC Jet Classification Dataset, or asks about evaluating this task. Reports classification accuracy.
This benchmark evaluates the ability of neural language models to judge the grammatical acceptability of Japanese sentences. It probes deep syntactic knowledge, particularly long-distance dependencies and linguistic phenomena, by measuring performance on both in-domain and out-of-domain acceptability judgments. Use when the user wants to benchmark on JCoLA, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).
Tests the accuracy and computational efficiency of a differentiable Material Point Method (MPM) simulator on geophysical flow benchmarks. It probes the framework's ability to reproduce free-surface dynamics, granular collapse rheology, and rigid-body contact against analytical or experimental ground truth, while measuring GPU acceleration speedups. Use when the user wants to benchmark on JAX-MPM Geophysical Benchmarks, or asks about evaluating this task. Reports normalized_runout.
Extractive machine reading comprehension in Japanese. It probes a model's ability to locate exact answer spans in Japanese Wikipedia text given a question, evaluating performance across different answer types, question reasoning types, and answer lengths. Use when the user wants to benchmark on JaQuAD, or asks about evaluating this task. Reports F1 score.
Evaluates Japanese sentence embeddings on domain-specific semantic textual similarity (STS) and information retrieval (IR) tasks. It probes the model's ability to capture fine-grained semantic similarity in clinical text and retrieve relevant question-answer pairs in an educational domain. Use when the user wants to benchmark on JACSTS, QABot, or asks about evaluating this task. Reports Spearman's rank correlation.
Evaluates large language models on Japanese financial domain knowledge across five distinct tasks: sentiment analysis, fundamental financial knowledge, CPA auditing, and two levels of financial planner exam questions. It probes the models' ability to understand and reason over domain-specific multiple-choice questions in Japanese. Use when the user wants to benchmark on chabsa, cma Basics, cpa Audit, fp2, security_sales_1, or asks about evaluating this task. Reports accuracy.
Evaluates open-ended legal reasoning capabilities of LLMs in the Japanese legal domain. It assesses their ability to generate structured, legally accurate arguments based on bar exam writing tasks. Use when the user wants to benchmark on Japanese Bar Exam Writing Task, or asks about evaluating this task. Reports expert_score.
Evaluates audio-language models on multi-track comparative reasoning by asking them to compare two music tracks and answer questions. It probes the model's ability to perform grounded, sentence-level comparative explanations versus simple binary or short-answer discrimination. Use when the user wants to benchmark on Jamendo-MT-QA, or asks about evaluating this task. Reports accuracy, LLM-as-a-Judge score.
Evaluates medical multimodal models on real-world diagnostic reasoning using clinical case images and questions. It probes both factual accuracy in close-ended QA and the model's ability to generate clinically sound reasoning across key points, inference steps, and evidence citation. Use when the user wants to benchmark on JAMA Clinical Challenge, or asks about evaluating this task. Reports Accuracy.
Evaluates automatic lyrics transcription (ALT) systems on readability-aware formatting, including punctuation, capitalization, line breaks, and background vocal annotations. It distinguishes errors by token type to measure how well models adhere to industry-standard musical semantics and prosodic structure. Use when the user wants to benchmark on Jam-ALT, Schubert Winterreise Dataset (SWD), or asks about evaluating this task. Reports WER.
Evaluates the robustness of large language models against adversarial jailbreaking attacks and defenses. It measures how effectively various attack methods can bypass safety filters (attack success rate) and how well defenses mitigate these attacks while maintaining normal functionality on benign prompts. Use when the user wants to benchmark on JBB-Behaviors, or asks about evaluating this task. Reports attack success rate (ASR).
Evaluates the robustness of LLM safety training against various jailbreak attacks by measuring the rate at which models produce harmful (BAD BOT), helpful (GOOD BOT), or ambiguous (UNCLEAR) responses to curated harmful prompts. Use when the user wants to benchmark on curated dataset, or asks about evaluating this task. Reports BAD BOT.