All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 12 installs2,582 views
Flm Audio EvalA

Evaluates native full-duplex audio-language models on speech understanding, speech generation, and real-time conversational capabilities. It measures how well models handle asynchronous text-audio streams, responsiveness to interruptions, and overall dialogue quality compared to specialized ASR/TTS systems and other full-duplex chatbots. Use when the user wants to benchmark on Fleurs-zh, LibriSpeech-clean, LlamaQuestions, Seed-TTS-en, Seed-TTS-zh, Custom Chinese Speech Instruction-Following S...

researchpythontesting
0
3
Flood Forecasting EvalA

Evaluates a model's ability to predict river water levels and forecast floods using spatiotemporal radar precipitation data. It probes the model's accuracy across multiple forecasting lead times (2h to 12h) and its robustness in capturing extreme hydrological events compared to baseline and deep learning models. Use when the user wants to benchmark on Goslar, Göttingen, or asks about evaluating this task. Reports NSE.

researchpythongo
0
3
FlopsA

This protocol evaluates the computational throughput and real-time efficiency of embedded CPU and GPU platforms. It measures peak floating-point operations per second (FLOPS) using a controlled matrix rotation kernel, and assesses practical system performance via end-to-end inference latency and power consumption on a robotic vision pipeline. Use when the user has predictions and gold and needs to compute FLOPS.

researchpythongo
0
3
Florence 2 EvalA

Evaluates a unified vision foundation model's zero-shot and fine-tuned capabilities across diverse computer vision tasks. It probes the model's ability to perform image captioning, visual question answering, object detection, referring expression comprehension, and semantic segmentation using a single sequence-to-sequence architecture. Use when the user wants to benchmark on COCO, Flickr30k, RefCOCO/+/g, VQAv2, ADE20K, or asks about evaluating this task. Reports CIDEr.

researchpythongo
0
3
Flores 200 EvalA

Evaluates machine translation quality across 200 languages by measuring meaning preservation and fluency. It compares automatic metrics (spBLEU, chrF++) against calibrated human judgments using the XSTS protocol, while also assessing translation safety/toxicity. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports XSTS.

researchpythongit
0
3
Flores101 Mt EvalA

Evaluates the translation quality of Neural Machine Translation (NMT) models trained on filtered pseudo-parallel corpora. It measures how well few-shot Quality Estimation (QE) based corpus filtering improves MT performance across low-resource and mid-resource language pairs compared to baselines and other filtering methods. Use when the user wants to benchmark on FLORES101, or asks about evaluating this task. Reports BLEU.

researchpythonperformance
0
3
Flores200 Mt EvalA

Evaluates multilingual machine translation quality across 60 languages and 234 translation directions. It specifically probes a model's ability to handle high-, medium-, and low-resource languages while mitigating directional degeneration in symmetric multi-way translation. Use when the user wants to benchmark on FLORES-200, or asks about evaluating this task. Reports COMET-22.

researchpythongo
0
3
Flow360 EvalA

Evaluates the accuracy of predicted optical flow fields on spherical 360° video frames, measuring both endpoint displacement and angular deviation. It also assesses egocentric activity recognition performance using these flow features to test rotation-invariant representation learning. Use when the user wants to benchmark on FLOW360, EGOK360, or asks about evaluating this task. Reports EPE.

researchpythongo
0
3
Flowbench EvalA

Evaluates computational workflow anomaly detection by benchmarking models on detecting injected CPU and HDD performance anomalies in distributed workflow execution logs. It tests the ability of tabular, graph, and text-based methods to identify anomalous nodes within Directed Acyclic Graph (DAG) workflow executions. Use when the user wants to benchmark on Flow-Bench, or asks about evaluating this task. Reports ROC-AUC.

businesspythongo
0
3
Flower Framework EvalA

Evaluates the scalability, heterogeneity handling, realism, and privacy overhead of the Flower federated learning framework across various datasets and device configurations. Use when the user wants to benchmark on Amazon Book Reviews, FEMNIST, RealWorld, CIFAR-10, FashionMNIST, or asks about evaluating this task. Reports training time.

businesspython
0
3
Flowtransformer EvalA

This evaluation protocol assesses the effectiveness of various transformer-based architectures for flow-based network intrusion detection. It systematically tests different input encodings, transformer blocks, and classification heads across three standard NIDS datasets to determine optimal configurations for accuracy, model size, and inference speed. Use when the user wants to benchmark on NSL-KDD, UNSW-NB15, CSE-CIC-IDS2018, or asks about evaluating this task. Reports F1 score.

researchpythongo
0
3
Flowxpert Mawi EvalA

Evaluates a network intrusion detection model's ability to classify benign versus malicious traffic flows in real-world IoT environments. It specifically probes robustness to severe class imbalance, feature sparsity mitigation via context-aware embeddings, and temporal generalization across different time periods. Use when the user wants to benchmark on MAWI, or asks about evaluating this task. Reports F1-Score.

researchpythongo
0
3
Flue EvalA

Probes French sequence classification capabilities across sentiment analysis, paraphrase identification, and natural language inference. Use when the user wants to benchmark on CLS, PAWSX, XNLI, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Fluidgym EvalA

Evaluates reinforcement learning algorithms for active flow control tasks, measuring their ability to stabilize fluid dynamics and reduce drag or enhance heat transfer. It probes algorithmic robustness, sample efficiency, and the capacity to transfer policies across dimensionalities and domain sizes. Use when the user wants to benchmark on FluidGym, or asks about evaluating this task. Reports mean reward per step.

researchpythongo
0
3
Fluidlab EvalA

Evaluates the ability of reinforcement learning and trajectory optimization algorithms to control complex, multi-phase fluid systems interacting with rigid bodies. It probes sample efficiency, gradient-based optimization stability, and sim-to-real transfer in high-dimensional, non-smooth fluid dynamics. Use when the user wants to benchmark on FluidLab, or asks about evaluating this task. Reports accumulated reward.

researchpythongo
0
3
Fluke Robustness EvalA

Evaluates how well NLP models maintain performance when subjected to minimal, linguistically-grounded perturbations (e.g., syntactic voice changes, negation, style shifts, geographical/temporal biases) across classification and generation tasks. It probes model brittleness to covariate shifts introduced by natural language modifications rather than adversarial noise. Use when the user wants to benchmark on KnowRef, Few-NERD, GSM8K, IFEval, or asks about evaluating this task. Reports Unrobustn...

researchpythongo
0
3
Flying Serving EvalA

Evaluates the runtime performance of an LLM serving engine under bursty, heterogeneous, and long-context workloads. It probes the system's ability to dynamically switch between data and tensor parallelism to optimize latency and throughput while maintaining memory efficiency compared to static and alternative dynamic baselines. Use when the user wants to benchmark on ShareGPT, CodeActInstruct, HumanEval, Synthetic Workloads, or asks about evaluating this task. Reports TTFT.

researchpythonperformance
0
3
Fma Genre Classification EvalA

Evaluates music information retrieval models on genre classification tasks using a large-scale, open music dataset. It probes the model's ability to map audio tracks to hierarchical genre labels (single-label or multi-label) using raw audio or precomputed features. Use when the user wants to benchmark on FMA (Free Music Archive), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Fnvls Bleu 1234A

Compute fnvls/bleu_1234 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of fnvls/bleu_1234.

developmentpython
0
3
Fnvls Bleu1234A

Compute fnvls/bleu1234 via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of fnvls/bleu1234.

developmentpython
0
3
Foca Malware Classification EvalA

This evaluation probes a model's ability to classify Android malware by fusing audio and visual representations derived from raw APK binaries. It measures supervised classification performance across multiple malware families and benign samples using standard accuracy and macro-F1 metrics. Use when the user wants to benchmark on CICMalDroid-2020, Mal-Net, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Focal EvalA

Evaluates end-to-end reasoning, error propagation, and user experience in multi-modal cascading agents. It assesses technical performance (ASR/TTS fidelity, tool calling) and behavioral quality (reasoning, semantic similarity, contextual consistency) across simulated customer service journeys. Use when the user wants to benchmark on FOCAL Customer Journeys, or asks about evaluating this task. Reports Accuracy Similarity.

ai-agentspythondatabase
0
3
Focustrack EvalA

Evaluates visual object tracking performance specifically for anti-UAV scenarios, probing a model's ability to maintain target localization under abrupt camera motion, extreme scale variations, and small target sizes in thermal infrared imagery. Use when the user wants to benchmark on AntiUAV, AntiUAV410, or asks about evaluating this task. Reports AUC.

researchpythongo
0
3
Fogmachine EvalA

Evaluates a discrete-event simulation framework that fuses dynamic scene graphs with urban environments to model hierarchical, interconnected spaces under partial observability. It probes the simulator's capacity to reproduce emergent temporal behaviors, the accuracy of state reconstruction when agent views are sparse, and the computational efficiency of the underlying simulation engine. Use when the user wants to benchmark on FOGMACHINE Scenarios (Bruchsal, Wenningstedt, Trier), or asks abou...

researchpythonnode
0
3
Foice Detection EvalA

Evaluates the ability of state-of-the-art audio deepfake detectors to distinguish real speech from face-to-voice (FOICE) synthesized speech, and assesses how fine-tuning on FOICE data affects robustness against unseen synthesis pipelines like SpeechT5. Use when the user wants to benchmark on FOICE, SpeechT5, or asks about evaluating this task. Reports EER.

researchpythonperformance
0
3
Folktables EvalA

Evaluates how fairness interventions affect predictive accuracy and fairness violations across different geographic regions and time periods. It probes the stability of fairness metrics under distribution shift and the efficacy of pre-processing, in-processing, and post-processing interventions on tabular demographic data. Use when the user wants to benchmark on Folktables (ACS PUMS), or asks about evaluating this task. Reports accuracy.

researchpythongit
0
3
Folktexts EvalA

Evaluates the calibration and predictive accuracy of language models when used as risk scorers for tabular prediction tasks. It probes whether models can accurately quantify outcome uncertainty (calibration) while maintaining discriminative power (AUC) on natural-language versions of tabular datasets. Use when the user wants to benchmark on folktexts, or asks about evaluating this task. Reports ECE.

researchpythongit
0
3
Followir EvalA

Evaluates whether information retrieval models can follow complex, long-form instructions derived from TREC narratives to determine document relevance. It probes the model's ability to interpret conditional, negated, and composite relevance criteria rather than relying solely on keyword matching. Use when the user wants to benchmark on Robust04, News21, Core17, or asks about evaluating this task. Reports p-MRR.

researchpythonperformance
0
3
Foodseg103 EvalA

Evaluates fine-grained semantic segmentation and ingredient localization in food images. It probes a model's ability to handle pixel-wise mask prediction under high appearance variability, long-tailed class distributions, and cross-domain generalization to unseen cuisines. Use when the user wants to benchmark on FoodSeg103, or asks about evaluating this task. Reports mIoU.

researchpythontesting
0
3
Footballdb EvalA

This benchmark evaluates the robustness and accuracy of Text-to-SQL systems when translating natural language questions into SQL queries across different database schema designs. It probes how data model complexity, training data size, and language model scale impact execution accuracy on real-world user queries. Use when the user wants to benchmark on FootballDB, or asks about evaluating this task. Reports exact execution matching (EX).

researchpythongo
0
3
Forb EvalA

Evaluates the quality of universal image embeddings for flat object retrieval across diverse 2D domains (e.g., logos, paintings, currency) under varying visual distortions. It probes both candidate rank accuracy and the matching score margin to assess out-of-distribution generalization. Use when the user wants to benchmark on FORB, or asks about evaluating this task. Reports mAP@5.

researchpythongo
0
3
Forc2025 EvalA

Evaluates hierarchical multi-label classification of academic papers into a taxonomy of 170 research fields. It tests zero-shot/few-shot prompting and weakly-labeled data integration for field-of-research prediction. Use when the user wants to benchmark on FoRC4CL 2025, or asks about evaluating this task. Reports Micro-F1.

researchpythongo
0
3
Force Prompting EvalA

Evaluates a video generation model's ability to adhere to specified physics-based force signals (local point forces or global wind fields) and produce visually realistic, physically plausible dynamics. It probes generalization across diverse objects, materials, and motion categories using human preference judgments. Use when the user wants to benchmark on Local Point Force Benchmark, Global Force Benchmark, or asks about evaluating this task. Reports 2AFC win rate.

ai-agentspythongo
0
3
Forensics Bench EvalA

This benchmark evaluates Large Vision Language Models (LVLMs) on their ability to detect and attribute image or video forgeries. It probes generalization and reasoning capabilities across five dimensions: semantics, modalities, tasks, forgery types, and generation models, using multi-choice visual questions. Use when the user wants to benchmark on Forensics-Bench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Foreseaqa EvalA

Evaluates a model's ability to perform temporally grounded, multimodal video understanding in surveillance settings. It probes precise temporal localization, identity-based search, and complex reasoning over long videos using text-only or image+text queries. Use when the user wants to benchmark on ForeSeaQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Foresight2 EvalA

Evaluates a fine-tuned LLM's ability to predict future biomedical concepts and clinical disorders from patient clinical timelines. It measures concept prediction accuracy via precision and recall across different temporal windows and candidate counts, and assesses clinical risk forecasting by checking how many of the top-5 predicted disorders match the ground truth for the next month. Use when the user wants to benchmark on MIMIC-III, or asks about evaluating this task. Reports Precision.

researchpythongo
0
3
Forest Change EvalA

Evaluates a vision-language agent's ability to perform joint change detection and natural language captioning on bi-temporal remote sensing imagery. It probes the model's capacity to segment deforestation and built-environment changes at the pixel level while generating accurate semantic descriptions of those changes. Use when the user wants to benchmark on Forest-Change, LEVIR-MCI-Trees, or asks about evaluating this task. Reports MIoU.

researchpythongit
0
3
Forge Manufacturing EvalA

Evaluates multimodal large language models on fine-grained manufacturing tasks, including workpiece verification, surface defect inspection, and assembly verification. It probes the models' ability to combine visual grounding with domain-specific knowledge to identify anomalies or classify conditions. Use when the user wants to benchmark on FORGE, or asks about evaluating this task. Reports exact-match accuracy.

researchpythongo
0
3
Forge Sid EvalA

Evaluates the quality of generated Semantic Identifiers (SIDs) for generative retrieval in industrial recommendation and search systems. It measures how well SIDs capture item relationships and distribute usage fairly, and assesses their impact on downstream retrieval hitrate and online transaction metrics. Use when the user wants to benchmark on FORGE, or asks about evaluating this task. Reports HR@K.

datapythongo
0
3
Forgery Localization EvalA

Evaluates pixel-level localization accuracy for detecting AI-generated and traditionally tampered image forgeries. It probes a model's ability to distinguish manipulated regions from authentic content by measuring spatial overlap and detection trade-offs against ground-truth masks. Use when the user wants to benchmark on OpenSDID, GIT10K, CocoGlide, Inpaint32K, IMD2020, NIST16, CASIA, or asks about evaluating this task. Reports F1-score.

researchpythongit
0
3
Forgotten Polygons EvalA

This benchmark probes the ability of multimodal large language models to recognize regular and irregular geometric shapes from images and accurately count their sides. It further evaluates multi-step visual-mathematical reasoning by requiring models to identify multiple shapes, map them to side counts, and compute their sum. Use when the user wants to benchmark on Forgotten Polygons, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Forkmerge EvalA

Evaluates the ability of auxiliary-task learning methods to mitigate negative transfer and improve target task performance across multi-task, multi-domain, and semi-supervised learning settings. It probes how well a model can dynamically combine or select auxiliary tasks without degrading the primary task. Use when the user wants to benchmark on NYUv2, DomainNet, AliExpress, CIFAR-10, SVHN, or asks about evaluating this task. Reports Δm.

researchpythongo
0
3
Formationeval EvalA

Evaluates large language models' domain knowledge in petroleum geoscience using a 505-question multiple-choice benchmark. It probes understanding across seven specialized subdomains, including petrophysics, reservoir engineering, and drilling, while measuring performance variance by model size, cost, and question difficulty. Use when the user wants to benchmark on FormationEval, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Formosanbench EvalA

Evaluates large language models and speech systems on three endangered Formosan Austronesian languages (Atayal, Amis, Paiwan) across machine translation, automatic speech recognition, and text summarization. It probes zero-shot, few-shot (10-shot), and fine-tuning adaptation capabilities in typologically complex, low-resource settings. Use when the user wants to benchmark on FormosanBench, or asks about evaluating this task. Reports BLEU.

researchpythongo
0
3
Formula Extraction EvalA

Evaluates the ability of PDF document parsers to accurately extract mathematical formulas and preserve their semantic meaning. It probes format variability handling, representational non-uniqueness, and semantic equivalence recognition beyond simple character matching. Use when the user wants to benchmark on PDF Formula Extraction Benchmark, or asks about evaluating this task. Reports LLM-as-a-Judge.

researchpythongit
0
3
Fortress EvalA

Evaluates LLM safeguard robustness against national security and public safety (NSPS) risks by measuring both the model's tendency to generate harmful content in response to adversarial prompts and its tendency to incorrectly refuse legitimate benign requests. Use when the user wants to benchmark on FORTRESS, or asks about evaluating this task. Reports Average Risk Score (ARS).

researchpythongo
0
3
Fotbcd EvalA

Evaluates cross-dataset generalization and geographic domain shift in building change detection models. It measures how well models trained on one geographic region or dataset perform when tested on entirely different datasets, highlighting the impact of training data diversity on remote sensing model transferability. Use when the user wants to benchmark on FOTBCD-Binary, LEVIR-CD+, WHU-CD, or asks about evaluating this task. Reports IoU.

researchpythongo
0
3
Foundationalecgnet Ecg EvalA

Evaluates a lightweight foundational model for ECG-based cardiac analysis, specifically testing its ability to classify signals as Normal/Abnormal and perform fine-grained disease classification across multiple cardiac conditions. Use when the user wants to benchmark on PTB-XL, CinC 2017, MedalCare-XL, PTB, or asks about evaluating this task. Reports F1-score.

researchpythongo
0
3
Four Gyre Rom EvalA

Evaluates the ability of a machine learning closure model to stabilize reduced-order models for turbulent geophysical fluid dynamics. Specifically, it probes whether an extreme learning machine can predict mode-dependent eddy viscosities to maintain long-time integration accuracy and statistical steady-state behavior in coarse-grained ocean circulation simulations. Use when the user wants to benchmark on Four-gyre barotropic circulation problem, or asks about evaluating this task. Reports L2-...

datapythontesting
0
3
Foura EvalA

Evaluates the capability of a Fourier-domain low-rank adapter to generate high-quality, diverse images for style transfer and concept editing. It also assesses the adapter's performance on standard language understanding benchmarks compared to baseline adapters like LoRA. Use when the user wants to benchmark on Paintings, Blue-Fire, 3D, Origami, GLUE, or asks about evaluating this task. Reports HPSv2.1.

researchpythongo
0
3