Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,849
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 7,729–7,752 of 20,849 skills

Emerge EvalA

Evaluates the ability of information extraction models to update knowledge graphs with emerging textual knowledge. It probes capabilities in extracting existing and new triples, linking emerging entities, and deprecating obsolete relations based on temporal text evidence. Use when the user wants to benchmark on EMERGE, or asks about evaluating this task. Reports recall.

researchpythongo
0
3
Embspatial Bench EvalA

Evaluates large vision-language models' ability to reason about egocentric spatial relations (e.g., above, below, left, right, close, far) within 3D embodied environments. It probes whether models can accurately localize objects and identify spatial configurations from a first-person perspective. Use when the user wants to benchmark on Embspacial-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Embrace 3k EvalA

Evaluates embodied vision-language models on step-wise, instruction-driven navigation and interaction in complex, partially observable environments. It probes three core capabilities: active exploration, dynamic spatial-semantic reasoning, and multi-stage goal execution under closed-loop perception-action constraints. Use when the user wants to benchmark on EmbRACE-3K, or asks about evaluating this task. Reports success.

researchpythongo
0
3
Embodyguard EvalA

Evaluates LLMs' physical safety and decision-making capabilities in embodied contexts by testing their ability to refuse unsafe instructions, interpret safety-critical goals, model state transitions, and sequence actions correctly. Use when the user wants to benchmark on EmbodyGuard, or asks about evaluating this task. Reports recall.

researchpythongo
0
3
Embodiedgpt Control EvalA

This evaluation probes the model's ability to generate executable sub-goal plans from visual inputs and translate them into low-level control actions in simulated robotic environments. It specifically tests closed-loop planning and few-shot policy adaptation across standard embodied AI benchmarks. Use when the user wants to benchmark on Franka Kitchen, Meta-World, or asks about evaluating this task. Reports success rate.

researchpythongo
0
3
Embodiedcomp EvalA

Evaluates how image compression codecs impact the performance of Vision-Language-Action (VLA) models in closed-loop robotic manipulation tasks under ultra-low bitrates. It measures whether compressed visual inputs cause task failure or require excessive inference steps, highlighting the disconnect between traditional visual fidelity metrics and embodied AI operational requirements. Use when the user wants to benchmark on EmbodiedComp, or asks about evaluating this task. Reports Success Rate (...

researchpythonperformance
0
3
Embodiedclaw EvalA

Evaluates the efficiency and executability of conversational workflow automation for embodied AI development tasks, including environment synthesis, trajectory collection, and VLA model evaluation. Use when the user wants to benchmark on RoboTwin, or asks about evaluating this task. Reports average task completion time.

researchpythonperformance
0
3
Embodiedbrain EvalA

Evaluates an embodied AI model's capabilities in general multimodal reasoning, 3D spatial perception, and long-horizon task planning across multiple public benchmarks and a custom simulation environment. Use when the user wants to benchmark on MM-IFEval, MMStar, MMMU, AI2D, OCRBench, BLINK, CV-Bench, EmbSpatial, ERQA, EgoPlan, EgoPlan2, EgoThink, Internal Planning, VLM-PlanSim-99, or asks about evaluating this task. Reports Action Pair Match F1-Score.

researchpythongo
0
3
Embodied Reasoner EvalA

Evaluates an agent's ability to perform long-horizon embodied interactive tasks by synergizing visual search, reasoning, and action. It probes spatial reasoning, self-reflection, and planning capabilities in both simulated and real-world environments. Use when the user wants to benchmark on Unspecified (Simulated & Real-world tasks), or asks about evaluating this task. Reports success_rate.

researchpythongo
0
3
Embodied Nav Safety EvalA

This protocol evaluates the safety and navigation performance of embodied agents against physical and model-based attacks. It measures task completion efficiency, path optimality, and goal satisfaction across diverse simulated and real-world environments. Use when the user wants to benchmark on Li et al. (2023), Kim et al. (2024), Khanna et al. (2024), Yin et al. (2024), Wang et al. (2024b), or asks about evaluating this task. Reports Success weighted by Path Length (SPL).

researchpythongo
0
3
Embodied Gov Bench EvalA

Evaluates embodied agent systems on runtime governance, safety, and accountability across seven dimensions: unauthorized capability invocation, runtime drift robustness, recovery success, policy portability, version upgrade safety, human override latency, and audit completeness. It shifts focus from pure task success to controllability and policy compliance under real-world perturbations. Use when the user wants to benchmark on EmbodiedGovBench, or asks about evaluating this task. Reports una...

researchpythonrust
0
3
Embodied Arena EvalA

This benchmark suite evaluates embodied AI models across perception, spatial reasoning, navigation, and task planning. It aggregates 22 diverse benchmarks to measure capabilities like 2D/3D question answering, instruction following in navigation, and complex task decomposition. Use when the user wants to benchmark on Embodied Arena, or asks about evaluating this task. Reports Exact Matching Accuracy.

researchpythongo
0
3
Embodied Ai Objectnav EvalA

Evaluates embodied AI agents' ability to navigate to target objects in 3D environments and perform manipulation tasks. It probes spatial reasoning, path efficiency, trajectory smoothness, and zero-shot generalization across different simulation domains and visual styles. Use when the user wants to benchmark on ProcTHOR-10k, ArchitecTHOR, AI2-iTHOR, RoboTHOR, ManipulaTHOR, Habitat 2022 ObjectNav, or asks about evaluating this task. Reports SR (Success Rate).

researchpythongo
0
3
Ember2024 EvalA

Evaluates malware classifiers on detection, family identification, and attribute prediction tasks across multiple file formats. It specifically probes robustness against concept drift and novel malware families using a temporally separated test set and a challenge set of evasive samples. Use when the user wants to benchmark on EMBER2024, or asks about evaluating this task. Reports ROC AUC.

researchpythonrust
0
3
Ember Malware Pipeline EvalA

Evaluates a multi-stage machine learning pipeline for detecting and classifying Windows PE files using static analysis features. It probes the model's ability to perform binary malware detection, hierarchical threat-type classification, family identification, and behavioral categorization. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Ember Malware Detection EvalA

Evaluates the performance and robustness of various machine learning classifiers for static malware detection on high-dimensional tabular features, specifically assessing the impact of dimensionality reduction techniques (PCA, LDA) on classification accuracy and discriminative power. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports accuracy.

researchpythonperformance
0
3
Embedding Clustering EvalA

This evaluation probes how well different embedding models capture underlying number-theoretic structures in numeric sequences. It measures the quality of the latent representation space by comparing clustering performance against ground-truth mathematical group labels versus unsupervised KMeans assignments. Use when the user wants to benchmark on number-theoretic-sequences, or asks about evaluating this task. Reports Silhouette Coefficient.

researchpythonperformance
0
3
Embedded Ai Companion EvalA

Evaluates the conversational quality, memory retrieval, personalization, and long-term memory extraction capabilities of an edge-deployed AI companion system over simulated multi-session interactions. Use when the user wants to benchmark on Synthetic User Simulation, or asks about evaluating this task. Reports Conversation Quality.

researchpythongo
0
3
Embbert Q EvalA

Evaluates tiny language models and baselines on resource-constrained embedded devices by measuring performance across classification and regression tasks under strict memory limits (≤2 MB). It probes the trade-off between model compression, hardware compatibility, and NLP task accuracy. Use when the user wants to benchmark on TinyNLP, GLUE, or asks about evaluating this task. Reports Accuracy, GLUE Average Score.

researchpythongo
0
3
Email Subject Line Generation EvalA

Evaluates a model's ability to generate highly abstractive, ultra-concise email subject lines from email body text. It probes extreme compression, informativeness, and fluency in a real-world email triaging context, distinguishing the task from standard text summarization. Use when the user wants to benchmark on AESLC, or asks about evaluating this task. Reports ROUGE-1.

researchpythongo
0
3
Elsa Ramd EvalA

Evaluates the robustness of Android malware detection models against feature-space and problem-space adversarial attacks, as well as temporal concept drift. It measures detection accuracy under increasing perturbation budgets while enforcing a strict false positive rate constraint on benign applications. Use when the user wants to benchmark on ELSA-RAMD Benchmark, or asks about evaluating this task. Reports TPR 100-FSA.

researchpythonrust
0
3
Elliptic Aml EvalA

Evaluates a graph neural network's ability to classify Bitcoin transactions as licit or illicit using structural and narrative features, while testing a retrieval-augmented generation pipeline for producing regulatory-aligned explanations. Use when the user wants to benchmark on Elliptic AML dataset, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Elasticc3 Clustering EvalA

Evaluates an unsupervised transfer learning method's ability to jointly cluster cells and genomic features across two datasets with differing distributions. It probes the model's capacity to elastically transfer clustering knowledge from an auxiliary dataset to a target dataset based on data similarity, without requiring labeled data or matching cluster counts. Use when the user wants to benchmark on Simulated scATAC-seq & scRNA-seq, Real data 1 (Human scRNA & scATAC), Real data 2 (Human & Mo...

researchpythonperformance
0
3
Elastic Scaling EvalA

Evaluates a deep learning job scheduler's ability to dynamically adjust GPU allocations and batch sizes to maximize cluster throughput and minimize job completion times. It probes how well the system handles compute-bound, communication-bound, and non-elastic workloads under varying job arrival patterns. Use when the user wants to benchmark on CIFAR100, Food101, or asks about evaluating this task. Reports SJS Efficiency.

researchpythongo
0
3