Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 5,833–5,856 of 23,891 skills
Evaluates a model's ability to reason about physical commonsense and intuitive physics by selecting the correct solution for a given goal from two options. It probes understanding of object affordances, material properties, and non-prototypical uses of everyday items. Use when the user wants to benchmark on PIQA, or asks about evaluating this task. Reports Accuracy.
Evaluates a model's ability to identify text spans corresponding to Patient, Intervention, and Outcome elements within clinical trial abstracts. Use when the user wants to benchmark on EBM-NLP, or asks about evaluating this task. Reports F-1.
Evaluates a model's ability to predict fine-grained hierarchical labels for tokens within identified P, I, and O spans. Use when the user wants to benchmark on EBM-NLP, or asks about evaluating this task. Reports F-1.
This evaluation probes a language model's ability to capture statistical patterns in diverse English text domains and its cross-domain generalization. It measures next-token prediction accuracy across academic, technical, legal, and conversational corpora. Use when the user wants to benchmark on The Pile, or asks about evaluating this task. Reports perplexity (BPB).
This benchmark probes the ability of NER and PII detection systems to accurately identify and classify personally identifiable information spans across highly heterogeneous, cross-domain text sources. It specifically evaluates cross-domain generalization and robustness to diverse, fine-grained PII entity types that are rarely seen together in standard training corpora. Use when the user wants to benchmark on PIIBench, or asks about evaluating this task. Reports span-level F1.
Evaluates the ability of PII masking models to correctly identify and classify sensitive information in text. It probes performance across diverse contexts, multilingual inputs, noisy formats, and evolving entity types. Use when the user wants to benchmark on PII Masking Dataset, or asks about evaluating this task. Reports non-identification.
Evaluates real-time performance and resource efficiency of TinyML models on embedded hardware by measuring inference latency, CPU/memory utilization, and prediction confidence across different platforms. Use when the user wants to benchmark on Gesture Classification, Keyword Spotting, MobileNet V2, or asks about evaluating this task. Reports Average Inference Latency (ms).
Evaluates the robustness of LLM prompt injection defenses against diverse attack strategies (heuristic, direct, adaptive, optimization-based) across multiple tasks and benchmarks. It measures the trade-off between maintaining legitimate task utility and preventing the execution of malicious injected instructions. Use when the user wants to benchmark on SQuAD v2, Dolly, NQ, InjecAgent, AgentDojo, AgentDyn, WASP, OPI, SEP, or asks about evaluating this task. Reports Attack Success Rate (ASR).
Probes multimodal physical reasoning by requiring models to interpret realistic visual scenarios, understand implicit physical conditions, and apply domain-specific knowledge across six physics domains. It evaluates both visual grounding and the ability to integrate symbolic reasoning with real-world constraints. Use when the user wants to benchmark on PhyX, or asks about evaluating this task. Reports accuracy.
Probes the ability of multimodal large language models to solve undergraduate-level physics problems that require integrating textual descriptions with complex diagrams. It specifically tests multi-step scientific reasoning, mathematical derivation, and conceptual understanding across eight distinct physics sub-disciplines. Use when the user wants to benchmark on PhysUniBench, or asks about evaluating this task. Reports accuracy.
Evaluates a model's ability to perform multilabel classification of cardiac abnormalities from 12-lead ECG recordings. It specifically probes robustness to severe class imbalance and the capacity to handle diverse lead configurations and sampling rates in long-sequence cardiac signals. Use when the user wants to benchmark on PhysioNet/CinC Challenge 2021, or asks about evaluating this task. Reports Macro AUPRC.
Evaluates the ability of models to forecast irregularly sampled multivariate time series generated from biological ordinary differential equations. It probes how well different architectures handle sparse, non-uniform temporal observations and varying levels of dynamical complexity. Use when the user wants to benchmark on Physiome-ODE, or asks about evaluating this task. Reports MSE.
Evaluates the utility of physically disentangled scene representations (geometry, albedo, lighting, camera) for downstream vision tasks. Probes robustness to out-of-distribution lighting and viewpoints, and measures how well learned features transfer to clustering, linear classification, and segmentation benchmarks. Use when the user wants to benchmark on CelebA, Buffy, BBT, RAF-DB, CelebA Mask, ShapeNet Cars, or asks about evaluating this task. Reports clustering_accuracy.
Evaluates a model's capacity for metric spatial reasoning, object enumeration, relative spatial comparisons, and topological/directional relationship understanding within real-world warehouse environments. Use when the user wants to benchmark on PhysicalAI-Spatial-Intelligence-Warehouse, or asks about evaluating this task. Reports normalized exact-match accuracy.
This benchmark evaluates a model's ability to reason about physical laws and detect physical commonsense violations in gameplay videos. It probes spatial, temporal, and meta-information-based physical reasoning through curated multi-choice questions. Use when the user wants to benchmark on PhysGame, or asks about evaluating this task. Reports accuracy.
Evaluates an agent's ability to reason about 2D Newtonian physics to solve goal-driven puzzles by placing dynamic objects. It probes sample-efficient learning and generalization across unseen task templates and action spaces. Use when the user wants to benchmark on PHYRE, or asks about evaluating this task. Reports AUCCESSION.
Evaluates a humanoid robot's ability to imitate human motion and follow pelvis trajectories using physically-grounded retargeting. It probes full-body tracking accuracy and partial-state path-following control across diverse locomotion categories on Unitree G1 and H1-2 robots. Use when the user wants to benchmark on PHUMA, Unseen Video, or asks about evaluating this task. Reports success_rate.
Evaluates phishing website detection models on temporally disjoint, real-world data with realistic base rates, while mitigating training-to-test leakage and varying difficulty levels. Use when the user wants to benchmark on PhreshPhish, or asks about evaluating this task. Reports Precision-Recall.
Classifies the sentiment polarity of specific phrases within tweets, requiring models to handle contextual disambiguation, sarcasm, and informal language. Use when the user wants to benchmark on Twitter2015-test, or asks about evaluating this task. Reports macro-averaged F1.
Evaluates the accuracy of predicting galaxy photometric redshifts from optical (grizy) photometric data across multiple redshift ranges. It probes a model's ability to minimize systematic bias, reduce catastrophic outliers, and maintain low error rates under varying data distributions. Use when the user wants to benchmark on Hyper Suprime-Cam Photometric Redshift Data, or asks about evaluating this task. Reports MAE.
Evaluates a 1D photochemical and climate model's ability to simulate atmospheric composition, photochemical networks, and radiative energy balance across diverse planetary environments. It probes whether the model can reproduce observed vertical gas profiles, cloud properties, and thermal structures without relying on unphysical surface fluxes. Use when the user wants to benchmark on Planetary Atmospheric Observations (Venus, Earth, Mars, Titan, Jupiter, WASP-39b), or asks about evaluating th...
This benchmark evaluates foundation models' ability to solve Olympiad-level physics problems using retrieval-augmented generation (RAG). It probes step-wise logical reasoning, correct application of physical laws, and robustness to noisy retrieved context. Performance is measured via an LLM-as-judge scoring framework that rewards both intermediate reasoning quality and final answer correctness. Use when the user wants to benchmark on PhoPile, or asks about evaluating this task. Reports Averag...
Evaluates automatic speech recognition performance on two low-resource, phonologically complex endangered languages (Archi and Kina Rutul) at the word, character, and phoneme levels. It specifically probes how training data frequency impacts phoneme recognition accuracy and error types. Use when the user wants to benchmark on Archi & Kina Rutul ASR, or asks about evaluating this task. Reports PER.
Evaluates the inference speed, energy efficiency, and classification accuracy of a GPU-accelerated binary neural network engine on mobile devices against standard frameworks. Use when the user has predictions and gold and needs to compute runtime.