
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates a model's ability to synthesize novel views of 3D scenes from a small set of input images using Gaussian splatting representations. It measures rendering quality, representation efficiency, and cross-dataset generalization. Use when the user wants to benchmark on RealEstate10K, ACID, or asks about evaluating this task. Reports PSNR.
Evaluates a system's ability to identify whether an incoming document contains novel information relative to a recent sliding window of previously seen documents in a text stream, using term specificity rather than pairwise similarity. Use when the user wants to benchmark on Real-world news stream, or asks about evaluating this task. Reports precision, recall, F1.
This benchmark evaluates multi-modal large language models on their ability to understand complex driving scenes using multi-view and multi-frame inputs. It probes three core capabilities: road environment perception, spatial relations recognition, and ego-centric reasoning through visual question answering. Use when the user wants to benchmark on NuPlanQA-Eval, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates no-reference image quality assessment (NR-IQA) models on their ability to predict human-perceived image quality without a pristine reference. It probes how well a model captures diverse authentic and synthetic distortions (e.g., blur, noise, exposure, haze) and maintains monotonic and linear correlation with crowd-sourced Mean Opinion Scores (MOS). Use when the user wants to benchmark on KonIQ-10k, LIVE Challenge, KADID-10k, TID2013, BIQ2021, IP102-IQA, or asks about ...
Evaluates recommendation algorithms for predicting drug-target and drug-disease interactions. It specifically probes the model's ability to handle bidirectional drug effects (therapeutic vs. adverse) and rank candidate pairs accurately under cold-start cross-validation scenarios. Use when the user wants to benchmark on Drug-Protein benchmark dataset, Drug-Disease benchmark dataset, or asks about evaluating this task. Reports AUPR.
Evaluates machine learning models for predicting Critical Heat Flux (CHF) in nuclear thermal-hydraulics, probing their ability to capture complex, multi-regime physical behaviors and produce well-calibrated, informative uncertainty estimates across different flow regimes. Use when the user wants to benchmark on NRC dataset, or asks about evaluating this task. Reports RMSPE.
Evaluates the quality of learned object-centric 3D representations across unsupervised segmentation, embodied object navigation, and relative depth ordering tasks. Use when the user wants to benchmark on ProcTHOR, RoboTHOR, CLEVR-3D, NYU Depth, or asks about evaluating this task. Reports ARI.
Evaluates network intrusion detection systems on imbalanced network traffic data by classifying records as normal or anomalous. It probes the model's ability to handle class imbalance and detect rare attack patterns in high-dimensional feature spaces. Use when the user wants to benchmark on NSL-KDD, or asks about evaluating this task. Reports F1 score.
Evaluates a federated transformer-based intrusion detection model on network traffic data to classify benign and malicious packets across five attack categories under a realistic class-imbalanced, distributed setting. Use when the user wants to benchmark on NSLKDD, or asks about evaluating this task. Reports detection performance.
Evaluates the ability of a hyperdimensional computing framework to detect and classify network intrusions in IoT environments. It probes the model's capacity to encode high-dimensional feature vectors, learn class prototypes, and accurately distinguish between normal traffic and specific attack types (DoS, probe, R2L, U2R). Use when the user wants to benchmark on NSL-KDD, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the ability of traditional time-series models and large language models to forecast half-hourly electricity prices in New South Wales, Australia. It specifically probes how well models integrate multimodal inputs (historical prices, weather, and market news) and tests for numerical reasoning capabilities, hallucination resistance, and strict output formatting compliance in a high-stakes forecasting task. Use when the user wants to benchmark on NSW-EPNews, or asks abou...
Evaluates the ability of audio generative models to reconstruct raw musical note waveforms and interpolate timbre and pitch in a learned latent space. Probes phase preservation, harmonic structure modeling, and the decoupling of pitch and timbre information. Use when the user wants to benchmark on NSynth, or asks about evaluating this task. Reports Classification accuracy.
Evaluates a model's ability to recognize human activities from RGB videos by leveraging skeleton-driven attention to focus on spatial-temporal regions of interest. It measures classification accuracy under standard cross-subject and cross-view protocols. Use when the user wants to benchmark on NTU-RGB+D, Northwestern-UCLA Multiview, or asks about evaluating this task. Reports accuracy.
Evaluates a learned neural metric's ability to correlate with human judgments of text generation quality. It probes semantic similarity, logical inference, and sentence likelihood capabilities across machine translation and image captioning domains. Use when the user wants to benchmark on WMT (Machine Translation), Flickr 8K, or asks about evaluating this task. Reports Pearson correlation.
Evaluates the ability of unbalanced optimal transport models to predict distributional shifts, mass creation (proliferation), and mass destruction (cell death) in heterogeneous populations under drug perturbation. Use when the user wants to benchmark on Synthetic Gaussian Mixture, Single-Cell Perturbation Response (Melanoma), or asks about evaluating this task. Reports weighted kernel MMD.
Evaluates a neural-network-based variational method for nuclear density functional theory by reproducing ground-state properties of finite nuclei and pasta phases, and benchmarking computational efficiency on GPU architectures. Use when the user wants to benchmark on Nuclear DFT Test Cases (Woods-Saxon, Finite Nuclei, Pasta Phases), or asks about evaluating this task. Reports binding energy.
Assesses large language models' deep scientific reasoning and domain-specific knowledge in nuclear physics, chemistry, and material science, without relying on provided context passages. It probes the model's ability to retrieve and apply expert-crafted factual and conceptual knowledge across multiple difficulty levels. Use when the user wants to benchmark on NuclearQA, or asks about evaluating this task. Reports accuracy.
Evaluates domain adaptive nuclei instance segmentation under cross-modality and cross-stain settings. It probes the model's ability to generalize to unseen cancer subdomains and different imaging modalities without target annotations. Use when the user wants to benchmark on BBBC039, Kumar, CPM17, DataSeg, or asks about evaluating this task. Reports Panoptic Quality (PQ).
Assesses a model's capability to predict various genomic features including histone markers, regulatory annotations, and splice sites. It tests fine-grained sequence understanding and multi-task classification across diverse genomic contexts. Use when the user wants to benchmark on Nucleotide Transformer Benchmark, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of models to perform instance-level segmentation of cell nuclei in H&E-stained histological images. It specifically probes robustness to tissue variability, high cell density, and ambiguous boundaries where manual annotation is difficult. Use when the user wants to benchmark on NuInsSeg, or asks about evaluating this task. Reports Dice score.
This benchmark probes fundamental numerical abilities in large language models, including number recognition, arithmetic operations, contextual retrieval, comparison, summarization, and logical reasoning. It evaluates how well models handle structured and unstructured numerical data across varying context lengths and noise levels. Use when the user wants to benchmark on NumericBench, or asks about evaluating this task. Reports accuracy.
Evaluates geospatial anomaly detection models on synthetic human mobility data. It probes the ability of algorithms to identify injected anomalous movement patterns across different granularities (staypoint, trip, agent) while controlling for demographic, temporal, and spatial factors. Use when the user wants to benchmark on NUMOSIM, or asks about evaluating this task. Reports Average Precision (AP).
Evaluates a driving planner's ability to navigate interactive scenarios in a closed-loop simulator. It measures success rates under both non-reactive (log-replay) and reactive (IDM/SMART) traffic conditions across routine validation splits and complex, human-curated scenarios. Use when the user wants to benchmark on Val14, Test14, interPlan, or asks about evaluating this task. Reports CLS-NR, CLS-R.
Evaluates autonomous driving planners in a closed-loop setting with realistic, reactive multi-agent traffic. It probes a planner's ability to handle complex, interactive driving scenarios by measuring overall success, robustness against catastrophic failures, and consistency across safety and comfort dimensions. Use when the user wants to benchmark on nuPlan-R, or asks about evaluating this task. Reports CLS.
Evaluates Vision-Language Models' ability to perform quantitative, agent-level risk assessment in autonomous driving. It probes spatio-temporal reasoning by testing whether models can predict collision risks, spatial distances, and temporal metrics based on visual sequences and optional physics-enhanced textual inputs. Use when the user wants to benchmark on NuRisk, or asks about evaluating this task. Reports MAE.
Evaluates sentiment classification and machine translation capabilities across 10 low-resource Indonesian local languages, Indonesian, and English. It probes cross-lingual transferability, multilingual training benefits, and data efficiency for underrepresented Austronesian languages. Use when the user wants to benchmark on NusaX, or asks about evaluating this task. Reports macro-F1.
Evaluates the capability of multi-view 3D object detection models to accurately localize and classify objects in autonomous driving scenes using camera inputs. It measures detection accuracy alongside computational efficiency and inference latency to assess real-time deployment feasibility. Use when the user wants to benchmark on NuScenes, or asks about evaluating this task. Reports NDS.
This benchmark evaluates the accuracy of end-to-end vectorized high-definition map construction from multi-camera images. It probes a model's ability to precisely predict instance-level road elements (lane dividers, pedestrian crossings, road boundaries) as continuous curves rather than rasterized masks or polylines. Use when the user wants to benchmark on NuScenes, or asks about evaluating this task. Reports mAP.
Assesses open-loop trajectory prediction accuracy of VLMs by measuring the distance between predicted and ground truth future paths at multiple time horizons. Use when the user wants to benchmark on nuScenes, or asks about evaluating this task. Reports L2 Error (m).
This benchmark evaluates a unified autonomous driving model across six core tasks: 3D object detection, multi-object tracking, BEV map segmentation, 3D occupancy prediction, motion prediction, and trajectory planning. It probes the model's ability to process heterogeneous multi-modal sensor data (LiDAR, multi-view cameras, and temporal sequences) and generate accurate, decoupled predictions for perception, prediction, and planning domains. Use when the user wants to benchmark on nuScenes, or ...
Evaluates novel view synthesis (NVS) methods on real-world handheld objects using only RGB inputs. It probes a model's ability to reconstruct 3D-consistent renderings from unconstrained, handheld camera trajectories that exhibit motion blur, occlusions, and pose estimation inaccuracies. Use when the user wants to benchmark on NVS-HO, or asks about evaluating this task. Reports PSNR.
This evaluation protocol assesses the capability of zero-shot text-to-speech models to synthesize nonverbal vocalizations (NVs) like breathing, laughter, coughing, and sighs alongside emotional speech. It measures speech intelligibility, speaker and emotion fidelity, acoustic quality, and the precise alignment of generated NVs with reference audio. Use when the user wants to benchmark on NVTTS, or asks about evaluating this task. Reports WER.
Evaluates the ability of image classification models to accurately categorize remote sensing scenes into one of 45 predefined land-use/land-cover categories. It probes robustness to realistic variations in spatial resolution, viewpoint, illumination, occlusion, and object pose that are common in aerial imagery. Use when the user wants to benchmark on NWPU-RESISC45, or asks about evaluating this task. Reports overall accuracy.
This benchmark evaluates a model's ability to classify breast cancer findings (benign vs. malignant vs. none) from high-resolution screening mammograms using only image-level labels. It also probes weakly supervised localization by measuring how well the model's generated saliency maps align with radiologist-annotated lesion segmentations. Use when the user wants to benchmark on NYU Breast Cancer Screening Dataset, or asks about evaluating this task. Reports AUC.
Evaluates weakly-supervised segmentation and classification performance on high-resolution mammography images for detecting malignant and benign breast lesions. Use when the user wants to benchmark on NYU Breast Cancer Screening Dataset v1.0, or asks about evaluating this task. Reports Dice similarity coefficient.
This benchmark evaluates models on academic graph mining tasks, including author disambiguation, scholar profiling, entity tagging, academic recommendation, question answering, paper source tracing, and influence prediction. It probes the ability of graph neural networks, retrieval systems, and LLMs to handle structured academic data, extract attributes from long texts, and perform ranking or prediction tasks on citation networks. Use when the user wants to benchmark on OAG-Bench, or asks abo...
Evaluates online sample selection methods for continual visual instruction tuning by measuring how effectively models learn from sequential data subsets. It probes the model's capacity for knowledge retention across tasks and its ability to adapt to new domains without catastrophic forgetting. Use when the user wants to benchmark on MICVIT, COAST, Adapt, Long Sequence, TRACE, or asks about evaluating this task. Reports $A_{last}$.
Evaluates joint extraction of aspect-based sentiment analysis elements (target, aspect, opinion, sentiment) at both sentence and review levels across multiple domains. Probes a model's ability to perform fine-grained, multi-element sentiment extraction and handle inter-sentence sentiment dynamics. Use when the user wants to benchmark on OATS, or asks about evaluating this task. Reports F1.
Probes a model's ability to localize and classify objects within images by generating bounding boxes and assigning confidence scores. It evaluates both proposal quality and final detection accuracy across varying scales, occlusions, and natural contexts. Use when the user wants to benchmark on PASCAL VOC 2007, MS COCO, or asks about evaluating this task. Reports mAP.
Evaluates object detection models trained on synthetic data against real-world baselines, probing their ability to generalize across domains without explicit domain adaptation. It tests how architectural choices (Transformers vs CNNs) and data augmentation strategies impact detection accuracy on geometric versus texture-heavy features. Use when the user wants to benchmark on DGTA-VisDrone, RarePlanes, Vehicle Detection, or asks about evaluating this task. Reports mAP@50.
This benchmark evaluates 6D object pose estimation for robotic manipulation tasks. It probes whether estimated poses are sufficiently accurate to enable successful physical assembly or grasping, rather than just measuring geometric alignment. Use when the user wants to benchmark on Industrial Object Pose Dataset, or asks about evaluating this task. Reports Average success probability.
Evaluates 3D grounding models' spatial reasoning and fine-grained object distinction capabilities by testing their ability to locate a target object among visually similar distractors in ensembled point cloud scenes. Use when the user wants to benchmark on OVE (ObjVariantEnsemble), or asks about evaluating this task. Reports ACC@0.25.
Evaluates the performance and scalability of Object-Oriented Databases (OODBs) under varying schema complexities and workload sizes. It specifically probes how database structure and clustering policies affect response times and I/O efficiency. Use when the user wants to benchmark on OCB (Object Clustering Benchmark), or asks about evaluating this task. Reports average response time.
Evaluates whether object-centric representations derived from zero-shot segmentation masks enable robust zero-shot classification under spurious background correlations, and compares them against slot-based OCL methods on unsupervised object discovery. Use when the user wants to benchmark on Movi-C, Movi-E, UrbanCars, ImageNet-D, ImageNet-9, Waterbirds, CounterAnimals, or asks about evaluating this task. Reports accuracy.
Evaluates the robustness of foundation segmentation models to synthetic surgical tool occlusions in endoscopic images. It probes whether models accurately segment visible tissue, avoid hallucinating into occluded regions, or maintain amodal completion under varying occlusion severities and prompt types. Use when the user wants to benchmark on CVC-300, CVC-ColonDB, ETIS, or asks about evaluating this task. Reports DSC.
Evaluates SAR foundation models on a suite of ocean observation tasks, including geophysical pattern classification, continuous regression for wave height and wind parameters, and iceberg object detection. It tests both zero-shot feature transferability and fine-tuning adaptability across diverse geophysical benchmarks. Use when the user wants to benchmark on Ocean Workbench, or asks about evaluating this task. Reports TenGeoP accuracy.
Compute Ochiroo/rouge_mn via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of Ochiroo/rouge_mn.
Evaluates an online continual learning model's ability to rapidly adapt to incoming data streams while retaining knowledge of past classes without catastrophic forgetting, under strict computational budgets and fixed feature extractors. Use when the user wants to benchmark on CGLM, CLOC, or asks about evaluating this task. Reports a_t.
Evaluates the performance of OCR systems on low-resource languages and scripts using both real and synthetically augmented PDF documents. It measures character-level accuracy to assess how OCR errors propagate and impact downstream tasks like machine translation. Use when the user wants to benchmark on OCR4MT, or asks about evaluating this task. Reports CER.
Evaluates Large Multimodal Models on five text-related visual tasks: text recognition, scene text-centric VQA, document-oriented VQA, key information extraction, and handwritten mathematical expression recognition. It probes the models' ability to perform precise visual pattern matching versus relying on semantic context, especially under challenging conditions like handwriting, multilingual text, blur, and complex layouts. Use when the user wants to benchmark on OCRBench, or asks about evalu...