
Claude Skills by qhjqhj00
github.com/qhjqhj00This benchmark evaluates advanced mathematical reasoning at the International Mathematical Olympiad level. It probes two distinct capabilities: deriving a unique integer answer from complex problem statements, and constructing step-by-step deductive proofs by solving decomposed sub-problems. Use when the user wants to benchmark on RIMO, or asks about evaluating this task. Reports exact-match grading.
Evaluates a model's ability to match 3D object patches and re-localize object instances in dynamically changing indoor environments. It measures feature matching robustness and 6DoF pose estimation accuracy under partial observations and contextual shifts. Use when the user wants to benchmark on 3RScan, or asks about evaluating this task. Reports Recall <0.1m, 10°.
Evaluates referring image segmentation on low-altitude drone imagery, probing the model's ability to accurately localize and segment referred objects despite challenges like category drift (tiny objects) and object drift (dense same-category scenes). Use when the user wants to benchmark on RIS-LAD, or asks about evaluating this task. Reports oIoU, mIoU.
This benchmark evaluates large language models' ability to perform medical logical reasoning and urological disease diagnosis. It probes the model's capacity to handle complex, real-world clinical scenarios involving subjective patient queries and multi-disease comorbidity reasoning. Use when the user wants to benchmark on RJUA-QA, or asks about evaluating this task. Reports F1 score (diagnosis & advice).
Evaluates LLMs' clinical capabilities across single-turn QA, diagnostic reasoning, and multi-turn dialogue using a urology-specific standardized patient dataset. It probes information gathering, diagnostic logic, treatment planning, and adherence to clinical workflows. Use when the user wants to benchmark on RJUA-SPs, or asks about evaluating this task. Reports Diagnosis Accuracy.
Evaluates the robustness and generalization of reinforcement learning-based dialogue management policies across varying simulated environments. It probes how well RL algorithms handle different domain sizes, user behavior profiles, and noisy speech input channels in task-oriented spoken dialogue systems. Use when the user wants to benchmark on PyDial simulated environments, or asks about evaluating this task. Reports average success rate.
Evaluates the sample efficiency and memory capabilities of reinforcement learning agents in partially observable environments. It probes whether a frozen language model can effectively compress historical observations to enable generalizable task solving without extensive finetuning. Use when the user wants to benchmark on RandomMaze, Minigrid (KeyCorridor), Procgen (Memory Mode), or asks about evaluating this task. Reports IQM of return.
Evaluates the trustworthiness (hallucination reduction) and helpfulness of multimodal large language models across generative, discriminative, and free-format tasks. Use when the user wants to benchmark on Object HalBench, MMHal-Bench, MHumanEval, AMBER, RefoMB, MMStar, or asks about evaluating this task. Reports response-level hallucination rate, trustworthiness win rate.
Evaluates a reinforcement learning framework for dynamic search ranking that adapts to evolving user intents over multiple search iterations using sequential feedback. It also benchmarks standard learning-to-rank performance on static datasets. Use when the user wants to benchmark on TREC 2016 Dynamic Domain, TREC 2017 Dynamic Domain, MQ2007, MQ2008, or asks about evaluating this task. Reports α-NDCG.
This evaluation probes the mathematical and out-of-domain reasoning capabilities of large language models trained with Reinforcement Learning with Verifiable Rewards (RLVR). It specifically tests how well entropy-aware credit assignment methods allocate learning signals across high-entropy tokens during chain-of-thought generation. Performance is measured by average accuracy and pass rate over multiple sampled reasoning paths. Use when the user wants to benchmark on AIME24, AIME25, AMC, MATH,...
Evaluates reward models' ability to correctly identify preferred responses based on substantive content rather than superficial stylistic cues. It probes sensitivity to subtle correctness differences, resistance to verbosity/style bias, and performance across diverse domains like math, code, and safety. Use when the user wants to benchmark on RM-Bench, or asks about evaluating this task. Reports Average Accuracy.
This benchmark evaluates robotic manipulation policies on memory-dependent, non-Markovian tasks. It probes a model's ability to retain and utilize historical visual and state information over long horizons to complete multi-step dual-arm manipulation sequences. Use when the user wants to benchmark on RMBench, or asks about evaluating this task. Reports success rate.
Evaluates a model's ability to score and rank candidate 3D RNA structural models by predicting their deviation from the true native structure. It probes the model's capacity to distinguish accurate conformations from decoys using only atomic coordinates and types. Use when the user wants to benchmark on RNA-Puzzles & FARFAR2 Decoys, or asks about evaluating this task. Reports RMSD.
Evaluates video-level road anomaly segmentation models in autonomous driving scenarios, specifically probing their ability to maintain prediction validity over time sequences and perform under real-time latency constraints. Use when the user wants to benchmark on Road Anomaly Segmentation Dataset, or asks about evaluating this task. Reports latency-aware metrics.
Evaluates graph neural networks and embedding methods for predicting traffic accident occurrences and counts on road network edges. It probes the models' ability to capture spatial-temporal dependencies, leverage graph structural features, and benefit from multitask or transfer learning across different U.S. states. Use when the user wants to benchmark on Traffic Accident Dataset, or asks about evaluating this task. Reports MAE.
Evaluates the ability to estimate traffic flow profiles for unsensed road segments by selecting similar roads based on topological embeddings or generating synthetic data. It probes how well graph-based feature similarity correlates with actual traffic pattern matching and the accuracy of generative models for sensorless traffic estimation. Use when the user wants to benchmark on Madrid Traffic Network, or asks about evaluating this task. Reports RMSE.
Evaluates vision-language models on visual question answering for Indian road scenes. It probes capabilities in object counting, object description, and surrounding scene description across diverse driving environments. Use when the user wants to benchmark on RoadscapesQA, or asks about evaluating this task. Reports exact-match accuracy.
This benchmark evaluates the quality of feature importance estimators in deep neural networks by measuring how model performance degrades when ranked important features are removed and the model is retrained. It probes whether an interpretability method correctly identifies pixels that the model actually relies on for prediction. Use when the user wants to benchmark on ImageNet, Birdsnap, Food 101, or asks about evaluating this task. Reports test accuracy.
Evaluates a model's ability to jointly detect aspects, sentiments, targets, and opinions across entire product/course reviews, capturing contextual dependencies that sentence-level methods miss. It probes review-level joint extraction for both triplet (aspect-sentiment-target) and quadruple (aspect-sentiment-target-opinion) formats. Use when the user wants to benchmark on ROAST Benchmark (Amazon_FF, Coursera, Hotels, Phones, Movies), or asks about evaluating this task. Reports F1 score.
Evaluates vision-language models' ability to perform single-step and multi-step spatial understanding and referring tasks in robotics contexts. It probes capabilities like 2D/3D relation reasoning, depth perception, and complex compositional spatial constraints in cluttered scenes. Use when the user wants to benchmark on CV-Bench, BLINK, RoboSpatial, RefSpatial-Bench, RefCOCO, or asks about evaluating this task. Reports Top-1 accuracy.
This benchmark evaluates the ranking accuracy and sample efficiency of generalist robot policies in real-world environments. It probes whether a distributed, pairwise comparison framework can reliably approximate an exhaustive oracle ranking across diverse scenes and tasks. Use when the user wants to benchmark on RoboArena Real-World Policy Evaluation, or asks about evaluating this task. Reports Pearson correlation r.
Evaluates multimodal large language models (MLLMs) as embodied brains in robotic manipulation. It probes five cognitive dimensions: instruction comprehension, perception reasoning, generalized planning, affordance prediction, and failure analysis across diverse real-world robotic tasks. Use when the user wants to benchmark on RoboBench, or asks about evaluating this task. Reports accuracy (%).
Evaluates the effectiveness of a large-scale bimanual manipulation dataset and a hierarchical annotation framework on Vision-Language-Action models across different robotic platforms. It probes the models' ability to generalize across task complexities, leverage multi-resolution annotations, and benefit from trajectory quality filtering. Use when the user wants to benchmark on RoboCOIN, or asks about evaluating this task. Reports success_rate.
Evaluates a robot's ability to generalize semantic knowledge by predicting object affordances, locations, and materials in unseen environments and inferring ranks of unseen semantic triples. It probes multi-relational embedding performance for common-sense reasoning in residential robotics. Use when the user wants to benchmark on AI2Thor, or asks about evaluating this task. Reports MRR.
This benchmark evaluates the out-of-distribution robustness of monocular depth estimation (MDE) models when exposed to real-world corruptions such as weather changes, sensor failures, and data processing anomalies. It measures how much model performance degrades relative to a clean baseline across multiple corruption types and severity levels. Use when the user wants to benchmark on KITTI-C, NYUDepth2-C, KITTI-S, or asks about evaluating this task. Reports mCE.
Evaluates the robustness and safety of vision-language models (VLMs) for end-to-end autonomous driving when subjected to real-world sensor corruptions (e.g., fog, rain, motion blur) and prompt corruptions (e.g., bit errors, malicious attacks). It probes the model's ability to maintain accurate trajectory prediction and low collision rates under degraded inputs. Use when the user wants to benchmark on RoboDriveBench, or asks about evaluating this task. Reports AvgL2.
Evaluates object detection models' ability to generalize across diverse, real-world, domain-specific visual tasks. It probes fine-tuning performance and zero-shot transfer capabilities on crowdsourced, practitioner-curated datasets spanning multiple imaging modalities. Use when the user wants to benchmark on Roboflow 100, or asks about evaluating this task. Reports mAP@.50.
Evaluates embodied reasoning capabilities of vision-language models on robotic manipulation tasks. It probes spatial understanding and generation (e.g., object grounding, grasp pose prediction) and temporal understanding and generation (e.g., motion trace reconstruction, multi-step planning) across diverse indoor and tabletop scenarios. Use when the user wants to benchmark on RoboInter-VQA, or asks about evaluating this task. Reports accuracy.
Evaluates a multimodal robot narration framework's ability to select key events, generate natural language summaries, and perform failure analysis (risk estimation, localization, explanation, recovery) on real-world household robot tasks. Use when the user wants to benchmark on RoboNar, or asks about evaluating this task. Reports Accuracy on failure analysis tasks.
Evaluates egocentric robot perception and navigation in crowded, unstructured environments. It probes multi-view 3D detection, 3D multi-object tracking, motion prediction, and 3D/BEV occupancy prediction using synchronized camera, LiDAR, and ultrasonic sensor data. Use when the user wants to benchmark on RoboSense, or asks about evaluating this task. Reports average precision.
Evaluates the ability of reinforcement learning algorithms to learn sequential robot manipulation tasks in simulation. It probes how well methods can handle randomized initial states, multi-stage objectives, and continuous control over fixed-horizon episodes. Use when the user wants to benchmark on robosuite, or asks about evaluating this task. Reports reward.
Evaluates how well vision foundation models support robot manipulation policies in simulation and real-world environments. It probes cross-modal spatial reasoning, task generalization across diverse manipulation suites, and robustness to sensor noise and platform differences. Use when the user wants to benchmark on LIBERO, MetaWorld, or asks about evaluating this task. Reports success rate.
Evaluates the ability of diffusion transformer policies to perform long-horizon robotic manipulation tasks across bi-manual, single-arm, and simulated environments. It probes stable training, observation tokenization, and generalization across different robot morphologies and action spaces. Use when the user wants to benchmark on Robotic Manipulation Task Suite, or asks about evaluating this task. Reports success_rate.
Evaluates vision-language models' ability to perform spatial understanding, metric measuring, 2D/3D referring, and multi-step visual tracing in cluttered environments. It probes geometric reasoning, depth estimation, and collision-free path planning for robotic manipulation. Use when the user wants to benchmark on CV-Bench, BLINK_val, RoboSpatial, Embspacial, Q-spatial, MSMU, Where2Place, RefSpatial-Bench, ShareRobot-Bench, VABench-V, TraceSpatial-Bench, RoboTwin, MMEtest, MMBenchdev, OK-VQA,...
Evaluates robot manipulation policies on a unified set of contact-rich pick-and-place and articulation tasks across multiple simulators. It probes both specialist and generalist vision-language-action models on success rates under standard and progressively challenging generalization levels. Use when the user wants to benchmark on ROBOVERSE Imitation Learning Benchmark, or asks about evaluating this task. Reports success rate.
Evaluates the robustness of monaural automatic speech recognition systems under noisy and reverberant conditions. It measures how well a decoupled frontend speech enhancement module improves the word error rate of a backend ASR model trained exclusively on clean speech. Use when the user wants to benchmark on WSJ0 SI-84, CHiME-2, LibriSpeech, or asks about evaluating this task. Reports WER.
Evaluates the robustness of end-to-end automatic speech recognition models against various stationary and non-stationary noise types at different signal-to-noise ratios (SNR). It also measures the degradation of recognition accuracy on clean speech when noise-adaptation techniques are applied. Use when the user wants to benchmark on Custom noisy speech dataset (7 noise types), or asks about evaluating this task. Reports WER.
Evaluates a model's ability to forecast future values in a financial time-series dataset given historical observations. It measures prediction accuracy using a capped relative error to prevent extreme outliers from dominating the score. Use when the user has predictions and gold and needs to compute robust mean relative error.
Evaluates multi-view depth estimation models on their ability to generalize across diverse domains, scales, and camera configurations. It specifically probes robustness to out-of-distribution cost volume statistics and tests absolute scale estimation without requiring scale alignment or depth range assumptions. Use when the user wants to benchmark on StaticThings3D, BlendedMVS, or asks about evaluating this task. Reports depth error metrics (e.g., RMSE, AbsRel, δ1).
Evaluates adversarial robustness of image classifiers under multi-norm threat models, testing whether pre-screening diagnostics (FOSC, RDI) reliably predict full attack performance and expose worst-case vulnerabilities masked by single-norm evaluations. Use when the user wants to benchmark on RobustBench CIFAR-10, or asks about evaluating this task. Reports robust_accuracy.
Evaluates the robustness of neural language models to non-adversarial character- and word-level input perturbations (e.g., typos, deletions, synonyms) while preserving semantic meaning. Use when the user wants to benchmark on TC, SA, NER, SS, QA (unspecified downstream datasets), or asks about evaluating this task. Reports accuracy.
Evaluates the cognitive alignment and visuocognitive reasoning of Vision-Language Models (VLMs) compared to humans on real-world, out-of-distribution autonomous driving scenarios from Peru. It probes how models and humans interpret complex, rare driving situations through open-ended, multiple-choice, and counterfactual/hypothetical visual question answering. Use when the user wants to benchmark on Robusto-1, or asks about evaluating this task. Reports Representational Similarity Analysis (RSA).
Evaluates the robustness of dense correspondence models (optical flow, scene flow, stereo) to 20 types of image corruptions by measuring the divergence between predictions on clean images and predictions on corrupted images. Use when the user wants to benchmark on Spring, or asks about evaluating this task. Reports R^c_EPE.
Compute the roc_auc_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute roc_auc_score, or asks how to score with roc_auc_score.
Compute the ROC metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute ROC, or asks how to score with ROC.
Evaluates the robustness of image-text matching models against adversarial perturbations injected into the retrieval gallery. It probes whether models rely on holistic semantic alignment or are easily misled by locally similar but semantically altered images and captions. Use when the user wants to benchmark on MS-COCO (RoCOCO variant), or asks about evaluating this task. Reports Recall@1.
Evaluates multimodal large language models on temporal segmentation and fine-grained behavioral annotation of rodent videos. It probes capabilities in long-video processing, distinguishing subtle or rare behaviors, and handling diverse experimental paradigms and camera angles. Use when the user wants to benchmark on Rodent-Bench-Long, Rodent-Bench-Short, or asks about evaluating this task. Reports Weighted Matthew’s Correlation Coefficient (MCC).
Evaluates a deep learning autoencoder's ability to reconstruct high-quality images from undersampled or noisy data across synthetic, MRI, and CT domains. It probes robustness to impulse noise, Fourier undersampling, and sparse tomographic projections compared to compressed sensing and standard autoencoders. Use when the user wants to benchmark on CIFAR-10, Cardiac Perfusion MRI, Larynx & Cardiac MRI, Speech MRI, ULB CT Dataset, or asks about evaluating this task. Reports NMSE.
This benchmark evaluates large language models' mathematical reasoning capabilities specifically in Romanian. It probes the ability to solve single-step and multi-step problems, handle verifiable numerical answers, and construct or verify mathematical proofs without relying on direct English translations. Use when the user wants to benchmark on RoMath, or asks about evaluating this task. Reports correctness.
This benchmark evaluates binary vulnerability detection capabilities on assembly language representations of C/C++ functions. It probes whether models can identify security flaws (e.g., buffer overflows, integer overflows) by analyzing machine code semantics and call graph context. Use when the user wants to benchmark on ROMEO, or asks about evaluating this task. Reports Accuracy.