Research
Research, evidence gathering, literature, reports, investigation, and synthesis
Browse research skills
Showing 7,729–7,752 of 20,849 skills
Evaluates the ability of information extraction models to update knowledge graphs with emerging textual knowledge. It probes capabilities in extracting existing and new triples, linking emerging entities, and deprecating obsolete relations based on temporal text evidence. Use when the user wants to benchmark on EMERGE, or asks about evaluating this task. Reports recall.
Evaluates large vision-language models' ability to reason about egocentric spatial relations (e.g., above, below, left, right, close, far) within 3D embodied environments. It probes whether models can accurately localize objects and identify spatial configurations from a first-person perspective. Use when the user wants to benchmark on Embspacial-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates embodied vision-language models on step-wise, instruction-driven navigation and interaction in complex, partially observable environments. It probes three core capabilities: active exploration, dynamic spatial-semantic reasoning, and multi-stage goal execution under closed-loop perception-action constraints. Use when the user wants to benchmark on EmbRACE-3K, or asks about evaluating this task. Reports success.
Evaluates LLMs' physical safety and decision-making capabilities in embodied contexts by testing their ability to refuse unsafe instructions, interpret safety-critical goals, model state transitions, and sequence actions correctly. Use when the user wants to benchmark on EmbodyGuard, or asks about evaluating this task. Reports recall.
This evaluation probes the model's ability to generate executable sub-goal plans from visual inputs and translate them into low-level control actions in simulated robotic environments. It specifically tests closed-loop planning and few-shot policy adaptation across standard embodied AI benchmarks. Use when the user wants to benchmark on Franka Kitchen, Meta-World, or asks about evaluating this task. Reports success rate.
Evaluates how image compression codecs impact the performance of Vision-Language-Action (VLA) models in closed-loop robotic manipulation tasks under ultra-low bitrates. It measures whether compressed visual inputs cause task failure or require excessive inference steps, highlighting the disconnect between traditional visual fidelity metrics and embodied AI operational requirements. Use when the user wants to benchmark on EmbodiedComp, or asks about evaluating this task. Reports Success Rate (...
Evaluates the efficiency and executability of conversational workflow automation for embodied AI development tasks, including environment synthesis, trajectory collection, and VLA model evaluation. Use when the user wants to benchmark on RoboTwin, or asks about evaluating this task. Reports average task completion time.
Evaluates an embodied AI model's capabilities in general multimodal reasoning, 3D spatial perception, and long-horizon task planning across multiple public benchmarks and a custom simulation environment. Use when the user wants to benchmark on MM-IFEval, MMStar, MMMU, AI2D, OCRBench, BLINK, CV-Bench, EmbSpatial, ERQA, EgoPlan, EgoPlan2, EgoThink, Internal Planning, VLM-PlanSim-99, or asks about evaluating this task. Reports Action Pair Match F1-Score.
Evaluates an agent's ability to perform long-horizon embodied interactive tasks by synergizing visual search, reasoning, and action. It probes spatial reasoning, self-reflection, and planning capabilities in both simulated and real-world environments. Use when the user wants to benchmark on Unspecified (Simulated & Real-world tasks), or asks about evaluating this task. Reports success_rate.
This protocol evaluates the safety and navigation performance of embodied agents against physical and model-based attacks. It measures task completion efficiency, path optimality, and goal satisfaction across diverse simulated and real-world environments. Use when the user wants to benchmark on Li et al. (2023), Kim et al. (2024), Khanna et al. (2024), Yin et al. (2024), Wang et al. (2024b), or asks about evaluating this task. Reports Success weighted by Path Length (SPL).
Evaluates embodied agent systems on runtime governance, safety, and accountability across seven dimensions: unauthorized capability invocation, runtime drift robustness, recovery success, policy portability, version upgrade safety, human override latency, and audit completeness. It shifts focus from pure task success to controllability and policy compliance under real-world perturbations. Use when the user wants to benchmark on EmbodiedGovBench, or asks about evaluating this task. Reports una...
This benchmark suite evaluates embodied AI models across perception, spatial reasoning, navigation, and task planning. It aggregates 22 diverse benchmarks to measure capabilities like 2D/3D question answering, instruction following in navigation, and complex task decomposition. Use when the user wants to benchmark on Embodied Arena, or asks about evaluating this task. Reports Exact Matching Accuracy.
Evaluates embodied AI agents' ability to navigate to target objects in 3D environments and perform manipulation tasks. It probes spatial reasoning, path efficiency, trajectory smoothness, and zero-shot generalization across different simulation domains and visual styles. Use when the user wants to benchmark on ProcTHOR-10k, ArchitecTHOR, AI2-iTHOR, RoboTHOR, ManipulaTHOR, Habitat 2022 ObjectNav, or asks about evaluating this task. Reports SR (Success Rate).
Evaluates malware classifiers on detection, family identification, and attribute prediction tasks across multiple file formats. It specifically probes robustness against concept drift and novel malware families using a temporally separated test set and a challenge set of evasive samples. Use when the user wants to benchmark on EMBER2024, or asks about evaluating this task. Reports ROC AUC.
Evaluates a multi-stage machine learning pipeline for detecting and classifying Windows PE files using static analysis features. It probes the model's ability to perform binary malware detection, hierarchical threat-type classification, family identification, and behavioral categorization. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports accuracy.
Evaluates the performance and robustness of various machine learning classifiers for static malware detection on high-dimensional tabular features, specifically assessing the impact of dimensionality reduction techniques (PCA, LDA) on classification accuracy and discriminative power. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports accuracy.
This evaluation probes how well different embedding models capture underlying number-theoretic structures in numeric sequences. It measures the quality of the latent representation space by comparing clustering performance against ground-truth mathematical group labels versus unsupervised KMeans assignments. Use when the user wants to benchmark on number-theoretic-sequences, or asks about evaluating this task. Reports Silhouette Coefficient.
Evaluates the conversational quality, memory retrieval, personalization, and long-term memory extraction capabilities of an edge-deployed AI companion system over simulated multi-session interactions. Use when the user wants to benchmark on Synthetic User Simulation, or asks about evaluating this task. Reports Conversation Quality.
Evaluates tiny language models and baselines on resource-constrained embedded devices by measuring performance across classification and regression tasks under strict memory limits (≤2 MB). It probes the trade-off between model compression, hardware compatibility, and NLP task accuracy. Use when the user wants to benchmark on TinyNLP, GLUE, or asks about evaluating this task. Reports Accuracy, GLUE Average Score.
Evaluates a model's ability to generate highly abstractive, ultra-concise email subject lines from email body text. It probes extreme compression, informativeness, and fluency in a real-world email triaging context, distinguishing the task from standard text summarization. Use when the user wants to benchmark on AESLC, or asks about evaluating this task. Reports ROUGE-1.
Evaluates the robustness of Android malware detection models against feature-space and problem-space adversarial attacks, as well as temporal concept drift. It measures detection accuracy under increasing perturbation budgets while enforcing a strict false positive rate constraint on benign applications. Use when the user wants to benchmark on ELSA-RAMD Benchmark, or asks about evaluating this task. Reports TPR 100-FSA.
Evaluates a graph neural network's ability to classify Bitcoin transactions as licit or illicit using structural and narrative features, while testing a retrieval-augmented generation pipeline for producing regulatory-aligned explanations. Use when the user wants to benchmark on Elliptic AML dataset, or asks about evaluating this task. Reports F1-score.
Evaluates an unsupervised transfer learning method's ability to jointly cluster cells and genomic features across two datasets with differing distributions. It probes the model's capacity to elastically transfer clustering knowledge from an auxiliary dataset to a target dataset based on data similarity, without requiring labeled data or matching cluster counts. Use when the user wants to benchmark on Simulated scATAC-seq & scRNA-seq, Real data 1 (Human scRNA & scATAC), Real data 2 (Human & Mo...
Evaluates a deep learning job scheduler's ability to dynamically adjust GPU allocations and batch sizes to maximize cluster throughput and minimize job completion times. It probes how well the system handles compute-bound, communication-bound, and non-elastic workloads under varying job arrival patterns. Use when the user wants to benchmark on CIFAR100, Food101, or asks about evaluating this task. Reports SJS Efficiency.