Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,827
skills in category
868
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,673–6,696 of 20,827 skills

Jrdb Par EvalA

Evaluates a model's ability to jointly recognize individual actions, social group activities, and global crowd-level activities in crowded panoramic scenes. It probes multi-granular activity recognition and hierarchical graph-based scene understanding. Use when the user wants to benchmark on JRDB-PAR, or asks about evaluating this task. Reports Overall F1 ($\mathcal{F}_a$).

researchpythongo
0
3
Jotto Game EvalA

Evaluates the quality of computed game-theoretic strategies in the word-guessing game Jotto by measuring expected guesses against a benchmark opponent and in self-play, as well as the equilibrium approximation error (epsilon). Use when the user wants to benchmark on Jotto (2-5 letter variants), or asks about evaluating this task. Reports epsilon.

researchpythongo
0
3
Jointdnn Benchmarks EvalA

Evaluates a layer-granular DNN offloading framework by measuring inference latency and energy consumption across standard discriminative, generative, and autoencoder neural network architectures on mobile-cloud setups. Use when the user wants to benchmark on JointDNN Deep Architecture Benchmarks (AlexNet, OverFeat, VGG16, Deep Speech, ResNet, NiN, Chair, Pix2Pix), or asks about evaluating this task. Reports latency.

researchpythonperformance
0
3
Joint Face Spoofing Forgery EvalA

Evaluates models' ability to detect physical face spoofing attacks and digital face forgeries using visual appearance and physiological rPPG cues. It measures cross-domain generalization and compares separate versus joint multi-task learning protocols. Use when the user wants to benchmark on SiW, 3DMAD, HKBU-MarsV2, MSU-MFSD, 3DMask, ROSE-Youtu, FaceForensics++, DFDC, CelebDFv2, or asks about evaluating this task. Reports AUC, EER.

researchpythontesting
0
3
John Muir Ant EvalA

Evaluates the ability of a genetic algorithm to evolve a navigation program for an artificial ant to traverse a complex, toroidal grid trail with gaps and high-difficulty sections. The benchmark measures how well the evolved program generalizes beyond the standard Santa Fe trail into a chaotic extended sector. Use when the user wants to benchmark on John Muir Ant Problem, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Jobbert EvalA

Evaluates a model's ability to normalize job titles by mapping them to standardized ESCO occupation labels using semantic similarity. It probes the model's capacity to handle hierarchical occupational taxonomies and filter out irrelevant contextual tokens like locations. Use when the user wants to benchmark on JobBERT Vacancy Titles, or asks about evaluating this task. Reports MRR.

researchpythongo
0
3
Jmmmu Pro EvalA

Evaluates multimodal language models' ability to perform integrated visual-textual reasoning on Japanese-language tasks where questions and reference images are combined into a single composite image. It specifically probes OCR capabilities, visual perception, and cross-modal alignment in a multilingual context. Use when the user wants to benchmark on JMMMU-Pro, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Jmmmu EvalA

This benchmark evaluates large multimodal models' ability to understand Japanese-language visual content and answer questions across multiple disciplines. It specifically probes the gap between general language translation capabilities (culture-agnostic subset) and deep cultural knowledge (culture-specific subset), revealing how models handle language variation bias and culturally grounded reasoning. Use when the user wants to benchmark on JMMMU, or asks about evaluating this task. Reports ac...

researchpythongo
0
3
Jmedbench EvalA

Evaluates Japanese biomedical large language models across five tasks: multiple-choice question answering, named entity recognition, machine translation, document classification, and semantic text similarity. It probes domain-specific knowledge, multilingual comprehension, and in-context learning capabilities. Use when the user wants to benchmark on JMedBench, or asks about evaluating this task. Reports F1-entity, Accuracy.

researchpythongo
0
3
Jigsaw Robot Manipulation EvalA

Evaluates spatial-temporal reasoning and hardware-agnostic robotic manipulation skills through a structured jigsaw puzzle assembly protocol. It measures vision-based segmentation, object recognition, pick planning success, and motion planning efficiency across three progressively complex physical tasks. Use when the user wants to benchmark on Jigsaw Manipulation Benchmark, or asks about evaluating this task. Reports Task Score.

researchpythongo
0
3
Jigsaw Puzzle Tiling EvalA

Evaluates object detection, recognition, and precise tiling assembly capabilities for modular robotic manipulation. It tests the pipeline's ability to segment oddly shaped pieces, recognize their identity, and assemble them with high spatial accuracy. Use when the user wants to benchmark on Jigsaw puzzle set, or asks about evaluating this task. Reports score.

researchpythongo
0
3
Jet Tagging Resilience EvalA

Evaluates the trade-off between classification performance (AUC) and model resilience (robustness to Monte Carlo simulation variations) in quark/gluon and top-quark jet tagging. It probes whether complex neural architectures generalize better to different physics simulators compared to simpler, physics-informed models. Use when the user wants to benchmark on Pythia 8 / Herwig 7 Jet Samples, or asks about evaluating this task. Reports AUC.

researchpythonangular
0
3
Jet Classification EvalA

Evaluates ultra-low-latency supervised classification of particle physics jet signatures on edge hardware. It probes the ability to distinguish rare boson/top-quark jets from common quark/gluon jets under strict microsecond latency and pipeline interval constraints. Use when the user wants to benchmark on LHC Jet Classification Dataset, or asks about evaluating this task. Reports classification accuracy.

researchpythongo
0
3
Jcola EvalA

This benchmark evaluates the ability of neural language models to judge the grammatical acceptability of Japanese sentences. It probes deep syntactic knowledge, particularly long-distance dependencies and linguistic phenomena, by measuring performance on both in-domain and out-of-domain acceptability judgments. Use when the user wants to benchmark on JCoLA, or asks about evaluating this task. Reports Matthews Correlation Coefficient (MCC).

researchpythongit
0
3
Jax Mpm Geophysical Benchmarks EvalA

Tests the accuracy and computational efficiency of a differentiable Material Point Method (MPM) simulator on geophysical flow benchmarks. It probes the framework's ability to reproduce free-surface dynamics, granular collapse rheology, and rigid-body contact against analytical or experimental ground truth, while measuring GPU acceleration speedups. Use when the user wants to benchmark on JAX-MPM Geophysical Benchmarks, or asks about evaluating this task. Reports normalized_runout.

researchpythonperformance
0
3
Jaquad EvalA

Extractive machine reading comprehension in Japanese. It probes a model's ability to locate exact answer spans in Japanese Wikipedia text given a question, evaluating performance across different answer types, question reasoning types, and answer lengths. Use when the user wants to benchmark on JaQuAD, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Japanese Sts Ir EvalA

Evaluates Japanese sentence embeddings on domain-specific semantic textual similarity (STS) and information retrieval (IR) tasks. It probes the model's ability to capture fine-grained semantic similarity in clinical text and retrieve relevant question-answer pairs in an educational domain. Use when the user wants to benchmark on JACSTS, QABot, or asks about evaluating this task. Reports Spearman's rank correlation.

researchpythongo
0
3
Japanese Financial Bench EvalA

Evaluates large language models on Japanese financial domain knowledge across five distinct tasks: sentiment analysis, fundamental financial knowledge, CPA auditing, and two levels of financial planner exam questions. It probes the models' ability to understand and reason over domain-specific multiple-choice questions in Japanese. Use when the user wants to benchmark on chabsa, cma Basics, cpa Audit, fp2, security_sales_1, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Japanese Bar Exam Legal Reasoning EvalA

Evaluates open-ended legal reasoning capabilities of LLMs in the Japanese legal domain. It assesses their ability to generate structured, legally accurate arguments based on bar exam writing tasks. Use when the user wants to benchmark on Japanese Bar Exam Writing Task, or asks about evaluating this task. Reports expert_score.

researchpythongo
0
3
Jamendo Mt Qa EvalA

Evaluates audio-language models on multi-track comparative reasoning by asking them to compare two music tracks and answer questions. It probes the model's ability to perform grounded, sentence-level comparative explanations versus simple binary or short-answer discrimination. Use when the user wants to benchmark on Jamendo-MT-QA, or asks about evaluating this task. Reports accuracy, LLM-as-a-Judge score.

researchpythongo
0
3
Jama Clinical Challenge EvalA

Evaluates medical multimodal models on real-world diagnostic reasoning using clinical case images and questions. It probes both factual accuracy in close-ended QA and the model's ability to generate clinically sound reasoning across key points, inference steps, and evidence citation. Use when the user wants to benchmark on JAMA Clinical Challenge, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Jam Alt EvalA

Evaluates automatic lyrics transcription (ALT) systems on readability-aware formatting, including punctuation, capitalization, line breaks, and background vocal annotations. It distinguishes errors by token type to measure how well models adhere to industry-standard musical semantics and prosodic structure. Use when the user wants to benchmark on Jam-ALT, Schubert Winterreise Dataset (SWD), or asks about evaluating this task. Reports WER.

researchpythonapi
0
3
Jailbreakbench EvalA

Evaluates the robustness of large language models against adversarial jailbreaking attacks and defenses. It measures how effectively various attack methods can bypass safety filters (attack success rate) and how well defenses mitigate these attacks while maintaining normal functionality on benign prompts. Use when the user wants to benchmark on JBB-Behaviors, or asks about evaluating this task. Reports attack success rate (ASR).

researchpythongit
0
3
Jailbreak EvalA

Evaluates the robustness of LLM safety training against various jailbreak attacks by measuring the rate at which models produce harmful (BAD BOT), helpful (GOOD BOT), or ambiguous (UNCLEAR) responses to curated harmful prompts. Use when the user wants to benchmark on curated dataset, or asks about evaluating this task. Reports BAD BOT.

researchpythongo
0
3