All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs1,982 views
SkillA

Open-weights chemistry reasoning model + verifiable reward functions from FutureHouse's ether0 (arXiv 2506.17238). Use to score model-generated chemistry outputs (SMILES validity, molecular completion, synthesis reasoning) against ground truth, or to run the open-weights ether0 model itself for chemistry reasoning. Also useful for visualizing molecules and reactions from SMILES.

ai-agentspythongo
0
3
2m Belebele EvalA

Multilingual reading comprehension across text and speech modalities. It probes a model's ability to understand spoken or written passages in 39 languages and answer multiple-choice questions based on them. Use when the user wants to benchmark on 2M-Belebele, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
360roam EvalA

Evaluates the capability of neural radiance field models to perform real-time, high-fidelity novel view synthesis on large-scale indoor scenes using 360° panoramic imagery. It probes the trade-off between rendering quality, computational efficiency, and geometric awareness in complex, unbounded indoor environments. Use when the user wants to benchmark on 360Roam Dataset, or asks about evaluating this task. Reports PSNR.

researchpython
0
3
3d 2d Vl Grounding EvalA

Evaluates a unified vision-language model's ability to ground natural language instructions to 3D objects and 2D regions, as well as answer 3D visual questions. It probes spatial reasoning, cross-modal alignment, and robustness to different 3D input representations (mesh-sampled vs. sensor RGB-D point clouds). Use when the user wants to benchmark on SR3D, NR3D, ScanRefer, RefCOCO, RefCOCO+, RefCOCOg, ScanQA, SQA3D, or asks about evaluating this task. Reports top-1 accuracy (Acc@25/50/75).

businesspythongo
0
3
3d Affordance Net EvalA

Evaluates 3D point cloud networks on visual object affordance understanding. It probes the model's ability to predict point-wise probabilistic scores for 18 affordance classes given full, partial, or rotated 3D shapes. Use when the user wants to benchmark on 3D AffordanceNet, or asks about evaluating this task. Reports mAP.

researchpythongo
0
3
3d Face Recon EvalA

Evaluates the accuracy and realism of 3D face reconstruction and generation from single 2D images. Probes shape reconstruction fidelity against ground truth meshes, texture identity preservation across novel poses, and the diversity of synthesized 3D faces. Use when the user wants to benchmark on NoW Benchmark, REALY 3D Benchmark, or asks about evaluating this task. Reports Per-vertex error (mm).

researchpythonexpress
0
3
3d Ids EvalA

Evaluates network intrusion detection systems on identifying malicious traffic flows in IoT and general network environments. It probes the model's ability to handle severe class imbalance, dynamic graph topologies, and both known and unknown attack patterns using binary and multi-class classification tasks. Use when the user wants to benchmark on CIC-ToN-IoT, CIC-BoT-IoT, EdgeIIoT, NF-UNSW-NB15-v2, NF-CSE-CIC-IDS2018-v2, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
3d Llm Benchmark EvalA

Evaluates whether Vision-Language Models (VLMs) and 3D LLMs genuinely understand 3D spatial reasoning or merely exploit 2D visual priors by rendering point clouds into images. It probes capabilities like object captioning, scene question-answering, and situation understanding across single-view, multi-view, and oracle-viewpoint settings. Use when the user wants to benchmark on 3D MM-Vet, ObjaverseXL-LVIS Caption, ScanQA, SQA3D, or asks about evaluating this task. Reports LLM-eval, EM.

researchpythongo
0
3
3d Medical Seg EvalA

Evaluates the segmentation performance of various 3D medical image architectures across multiple public datasets. It probes whether newer architectures genuinely outperform established U-Net baselines when trained under standardized, hardware-scaled conditions without external advantages like ensembling or pretraining. Use when the user wants to benchmark on BTCV, ACDC, LiTS, BraTS, KiTS, AMOS, or asks about evaluating this task. Reports DSC score [%].

researchpythongo
0
3
3d Mir EvalA

Probes the ability of medical imaging models to retrieve relevant 3D CT volumes based on lesion characteristics. It evaluates retrieval accuracy for binary lesion presence (flag) and morphological size categories (group) across four anatomical regions. Use when the user wants to benchmark on 3D-MIR, or asks about evaluating this task. Reports Average Precision (AP).

researchpythongo
0
3
3d Multimodal EvalA

Evaluates a 3D large multimodal model's ability to understand spatial scenes and generate accurate text responses. It probes free-form question answering about 3D environments and object-centric dense captioning grounded in 3D coordinates. Use when the user wants to benchmark on ScanQA, SQA3D, ScanRefer, Nr3D, or asks about evaluating this task. Reports CiDEr.

researchpythonperformance
0
3
3d Noc Power Thermal Reliability EvalA

Evaluates a simulation platform for predicting power consumption, thermal distribution, and reliability (MTTF) of 3D Networks-on-Chip under synthetic and application workloads. It compares TSV-based 3D-NoC designs against monolithic and 2D-IC alternatives, and assesses the impact of different floorplans and cooling strategies on thermal stress and failure rates. Use when the user wants to benchmark on PARSEC benchmark suite, Synthetic benchmarks (Matrix, HotSpot, Uniform, Transpose), or asks ...

researchpython
0
3
3d Obj Det Seg EvalA

Evaluates a model's ability to detect and segment 3D objects in indoor scenes using point cloud inputs. It probes spatial reasoning and instance-level understanding by measuring how well the model generalizes from synthetic internet-scale data to real-world scanned environments. Use when the user wants to benchmark on ScanNet, SceneVerse++, or asks about evaluating this task. Reports AP.

researchpythongo
0
3
3d Pose Estimation Manifold EvalA

Evaluates a model's ability to estimate 3D joint poses of articulated objects (mice, fish, human hands) from depth images. The benchmark probes continuous structured prediction on Lie group manifolds, requiring the model to output kinematic chain or tree configurations that align with ground-truth skeletal models. Use when the user wants to benchmark on Mouse, Fish, Human hand, or asks about evaluating this task. Reports average joint error.

researchpythonperformance
0
3
3d Scene Understanding EvalA

Evaluates a 3D vision-language model's ability to perform visual grounding, dense captioning, and situated question answering on indoor RGB-D scenes. It probes the model's capacity for precise object referencing, spatial reasoning, and open-ended language generation conditioned on 3D scene context. Use when the user wants to benchmark on ScanRefer, Multi3DRefer, Scan2Cap, ScanQA, SQA3D, or asks about evaluating this task. Reports Acc@0.5.

researchpython
0
3
3d Segmentation EvalA

Evaluates a unified transformer-based model's ability to perform instance and semantic segmentation on 3D point clouds derived from raw RGB-D sensor data. It probes cross-modal feature fusion between 2D images and 3D coordinates, and tests robustness to real-world sensor noise and misalignments compared to mesh-sampled inputs. Use when the user wants to benchmark on ScanNet, ScanNet200, or asks about evaluating this task. Reports mAP, mIoU.

researchpythongo
0
3
3d Semantic Segmentation EvalA

This benchmark evaluates a model's ability to perform 3D semantic segmentation on indoor scenes. It probes the model's capacity to assign per-point semantic labels to 3D point clouds and assesses performance across head, common, and long-tail object categories. Use when the user wants to benchmark on ScanNet/ScanNet200, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
3d Shape Retrieval EvalA

Evaluates algorithms for retrieving geometrically and topologically similar 3D shapes from a database given a query shape. Probes pose invariance, shape descriptor robustness, and retrieval ranking accuracy. Use when the user wants to benchmark on NIST shape benchmark, or asks about evaluating this task. Reports precision-recall.

researchpythongo
0
3
3d Spatial Reasoning EvalA

Evaluates a model's ability to perform 3D visual grounding and situated question answering by reasoning over object coordinates and spatial relations in 3D scenes. It probes whether the model can accurately locate objects based on natural language instructions and answer spatial questions about scene layouts without linguistic interference. Use when the user wants to benchmark on ScanRefer, Multi3DRef, SQA3D, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
3d Spatial Vqa EvalA

Tests a model's capacity for 3D spatial reasoning and scene understanding by answering questions about object counts, distances, directions, and room sizes. It evaluates how well foundation models can leverage automatically generated scene graphs and point cloud data for grounded visual question answering. Use when the user wants to benchmark on SceneVerse++ VQA, or asks about evaluating this task. Reports MCA Accuracy.

researchpythongo
0
3
3d Visual Grounding EvalA

Evaluates a model's ability to localize a specific 3D object within a scene based on a natural language description. It tests multimodal fusion of 3D point clouds, synthetic 2D views, and language to perform object classification and referring. Use when the user wants to benchmark on Nr3D, Sr3D, ScanRefer, or asks about evaluating this task. Reports referring accuracy.

researchpythongo
0
3
3d Vln EvalA

Evaluates a model's ability to navigate 3D environments based on natural language instructions. It measures path efficiency, success in reaching targets, and robustness to scale calibration and data distribution shifts. Use when the user wants to benchmark on R2R, NaVILA, SceneVerse++ VLN, or asks about evaluating this task. Reports SR.

researchpythongo
0
3
3d Vqa EvalA

Evaluates a model's ability to answer natural language questions about 3D indoor scenes using only multi-view RGB images. It probes spatial reasoning, semantic understanding, and zero-shot generalization across different embodied agent scenarios. Use when the user wants to benchmark on ScanQA, SQA3D, MSR3D, or asks about evaluating this task. Reports EM@1.

researchpythongo
0
3
3dmem Bench EvalA

Evaluates an embodied 3D agent's ability to manage long-term spatial-temporal memory and execute complex, multi-room tasks. It probes the model's capacity for in-domain generalization, in-the-wild robustness, and long-horizon reasoning across navigation, question answering, and scene captioning. Use when the user wants to benchmark on 3DMem-Bench, or asks about evaluating this task. Reports success rate (SR).

researchpythongo
0
3
3doc Bench EvalA

This benchmark evaluates a model's ability to generate text-to-image outputs that strictly adhere to 3D layout constraints, handle complex inter-object occlusions, and maintain correct object orientations and visibility orders. It probes depth-consistent scene composition, attribute binding to specific objects, and overall image fidelity under varying camera viewpoints. Use when the user wants to benchmark on 3DOc-Bench, or asks about evaluating this task. Reports depth ordering.

researchpythongo
0
3
3dses EvalA

Semantic segmentation of indoor Terrestrial Laser Scanning (TLS) point clouds. It probes a model's ability to classify 3D points into semantic categories (e.g., furniture, structural elements, clutter) using geometric coordinates and optionally Lidar intensity features. Use when the user wants to benchmark on 3DSES, or asks about evaluating this task. Reports mIoU.

researchpythongo
0
3
3mdbench EvalA

Evaluates Large Vision-Language Models in realistic telemedicine consultations by simulating multi-agent dialogues between a doctor and a temperament-based patient. It probes diagnostic accuracy from multimodal inputs (images + text) and assesses clinical competence and dialogue quality. Use when the user wants to benchmark on 3MDBench, or asks about evaluating this task. Reports F1 Score.

researchpythongo
0
3
4seasons EvalA

Evaluates visual SLAM and long-term localization for autonomous driving under challenging cross-season, multi-weather, and long-term environmental changes. Specifically probes visual odometry, global place recognition, and map-based visual localization capabilities. Use when the user wants to benchmark on 4Seasons, or asks about evaluating this task. Reports horizontal RMSE.

researchpythonperformance
0
3
5g Madrl Sumrate EvalA

Evaluates the ability of a multi-agent deep reinforcement learning framework to optimize the 3D placement and trajectory of mobile access points in dynamic 5G networks, balancing sum-rate maximization against user mobility and interference. Use when the user wants to benchmark on Custom 5G Network Simulation, or asks about evaluating this task. Reports sum-rate.

researchpythongo
0
3
6dof Camera Tracking EvalA

Evaluates the tracking accuracy of a 6-DoF autonomous camera algorithm in a simulated surgical environment. It also measures how different camera control strategies impact human rater accuracy when assessing surgical skill from video. Use when the user wants to benchmark on da Vinci wire chaser simulation, or asks about evaluating this task. Reports assessment_error.

researchpythongo
0
3
6dof Visual Localization EvalA

Evaluates a model's ability to estimate 6-degree-of-freedom camera poses for query images against a reference 3D model, specifically testing robustness to drastic changes in lighting (day/night), weather, and seasonal vegetation. Use when the user wants to benchmark on Aachen Day-Night, RobotCar Seasons, CMU Seasons, or asks about evaluating this task. Reports translation error and rotation error.

researchpythontesting
0
3
APESsrcA

Evaluates the faithfulness of abstractive summaries by verifying if factual claims (masked as cloze questions) in the reference summary can be correctly answered using only the generated summary, compared against a gold-standard answer derived from the source context. Use when the user has predictions and gold and needs to compute APESsrc.

researchpythongo
0
3
AUCA

This evaluation probes the capability of machine learning classifiers to distinguish signal events from background noise in high-energy particle physics simulations. It measures how well different algorithms and feature sets rank positive (signal) instances higher than negative (background) ones across varying data statistics. Use when the user has predictions and gold and needs to compute AUC.

datapythongo
0
3
CIDErA

Evaluates how well automatic image description metrics correlate with human consensus on sentence similarity. It probes the ability of generated or candidate sentences to accurately describe an image by measuring alignment with multiple human-generated reference descriptions. Use when the user has predictions and gold and needs to compute CIDEr.

documentationpythongo
0
3
EASA

This evaluation validates the Emotional Attitude Score (EAS) metric by measuring its consistency with human judgment on word-level sentiment polarity. It probes whether the metric's pseudo-log-likelihood-based scores reliably capture positive or negative emotional attitudes toward ambiguous attitude words in gender-inclusive contexts. Use when the user has predictions and gold and needs to compute EAS.

researchpythongo
0
3
MDBI MTBIA

Evaluates the robustness and safety of autonomous driving systems by quantifying how far and how long the vehicle operates between human interventions (disengagements). It enables unbiased comparison across different AV platforms and road environments by normalizing disengagement frequency with spatial and temporal data. Use when the user has predictions and gold and needs to compute MDBI, MTBI.

devopspythongo
0
3
MENLIA

Evaluates the robustness and alignment with human judgment of reference-based and reference-free evaluation metrics for machine translation and summarization, particularly under adversarial conditions. Use when the user has predictions and gold and needs to compute Pearson correlation.

researchpythongo
0
3
NDCG@10A

Evaluates how well internal model representations (hidden states) can predict token-level information importance in summarization tasks. It probes whether specific transformer layers or cross-layer combinations encode salience distributions consistent with empirical importance derived from summary persistence. Use when the user has predictions and gold and needs to compute NDCG@10.

researchpythongo
0
3
PSNRA

Evaluates the trade-off between file size reduction and image fidelity when encoding radio astronomy data using JPEG2000. It benchmarks both lossless and lossy compression modes to determine the compression ratio at which visual artifacts first appear. Use when the user has predictions and gold and needs to compute PSNR.

researchpythongo
0
3
SDRA

Evaluates how effectively audio separation models isolate specific stems from a mixed recording while preserving signal integrity. It quantifies the ratio of target source energy to residual interference and distortion energy. Use when the user has predictions and gold and needs to compute SDR.

researchpythongo
0
3
TCAV ScoreA

Measures the influence of human-defined emotional concepts (physiognomy, utterance polarity, voice pitch) on a multimodal emotion recognition model's decisions using Concept Activation Vectors. It quantifies how much each concept drives the model's classification decisions across different network layers. Use when the user has predictions and gold and needs to compute TCAV score.

datapythongo
0
3
TCTBA

Evaluates the throughput and resource allocation efficiency of RIS-aided mobile edge computing systems by measuring the total computation task bits successfully completed under varying network conditions. Use when the user has predictions and gold and needs to compute TCTB.

researchpythongo
0
3
TECA

Evaluates the joint optimization of AI service placement and resource allocation in mobile edge computing by measuring the trade-off between computation time and energy consumption across varying network scales and task characteristics. Use when the user has predictions and gold and needs to compute TEC.

researchpythongo
0
3
TTSDSA

Measures the distributional distance between synthetic and real speech across five key factors: environment, speaker identity, prosody, intelligibility, and general speech distribution. It evaluates TTS system quality without relying on subjective Mean Opinion Scores (MOS) or simple mean-based metrics. Use when the user has predictions and gold and needs to compute TTSDS.

researchpythongo
0
3
A2seek EvalA

Evaluates multimodal models' ability to detect, localize, and semantically reason about anomalies in aerial drone-view videos. It probes spatial grounding accuracy, temporal anomaly detection, and the generation of contextually grounded natural language explanations. Use when the user wants to benchmark on A2Seek, or asks about evaluating this task. Reports AP_c, mIoU.

researchpythongo
0
3
A3 EvalA

Evaluates mobile GUI agents on completing multi-step tasks across 20 real-world Android applications. It probes both final task completion capability and the agent's ability to navigate intermediate essential states without getting stuck or making terminal errors. Use when the user wants to benchmark on A3, or asks about evaluating this task. Reports Task Success Rate (SR).

researchpythongo
0
3
Aa Omniscience EvalA

Evaluates large language models' factual recall and knowledge calibration across domain-specific questions. It measures how reliably models provide correct answers versus hallucinating or abstaining when uncertain, highlighting the gap between raw accuracy and factual reliability. Use when the user wants to benchmark on AA-Omniscience, or asks about evaluating this task. Reports Omniscience Index.

researchpythongo
0
3
Abc EvalA

This benchmark evaluates large language models' ability to understand symbolic music and follow instructions using text-based ABC notation. It probes capabilities ranging from basic syntax parsing and error detection to segment-level reasoning and sequence-level musical analysis like genre or emotion recognition. Use when the user wants to benchmark on ABC-Eval, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Abcfair EvalA

Evaluates the trade-off between predictive performance and fairness across diverse real-world settings. It probes how different intervention stages, sensitive feature compositions, fairness notions, and output distributions impact a model's ability to satisfy fairness constraints while maintaining accuracy. Use when the user wants to benchmark on SchoolPerformance, ACSPublicCoverage, or asks about evaluating this task. Reports AUROC.

researchpythonperformance
0
3
Abdomenct1k EvalA

This benchmark evaluates the ability of 3D medical image segmentation models to accurately delineate abdominal organs (liver, kidney, spleen, pancreas) under clinically challenging conditions. It specifically probes generalization across unseen medical centers, CT contrast phases, and severe pathologies like tumors, while measuring both volumetric overlap and boundary precision. Use when the user wants to benchmark on AbdomenCT-1K, or asks about evaluating this task. Reports DSC.

researchpythontesting
0
3