All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,053 views
Cell Instance Segmentation EvalA

This evaluation probes a model's ability to perform precise instance segmentation on challenging biomedical microscopy images. It specifically tests handling of overlapping, irregularly shaped cells across varying contrast modalities (brightfield, phase-contrast, fluorescence) and object densities. Use when the user wants to benchmark on LIVECell, EVICAN2, ISBI2014, Revvity-25, or asks about evaluating this task. Reports AP (Average Precision).

researchpythongo
0
3
Cello EvalA

This benchmark evaluates the causal reasoning capabilities of Large Vision-Language Models (LVLMs) across four levels of the Ladder of Causation: discovery, association, intervention, and counterfactual. It probes whether models can correctly identify causal relationships, handle confounding and collider biases, reason about interventions, and generate counterfactual explanations based on visual scenes. Use when the user wants to benchmark on CELLO, or asks about evaluating this task. Reports...

researchpythongo
0
3
Cendol EvalA

This evaluation protocol assesses the language proficiency, generalization capability, and local cultural commonsense reasoning of instruction-tuned LLMs across Indonesian and nine indigenous languages. It probes zero-shot performance on seen and unseen tasks and languages, as well as nuanced understanding of regional proverbs, figures of speech, and story endings. Use when the user wants to benchmark on COPAL-ID, MABL, IndoStoryCloze, MAPS, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Certified Malware Detection EvalA

Evaluates the robustness of malware classifiers against metamorphic evasion attacks and synthetic feature-space perturbations. It probes whether randomized smoothing and majority voting can maintain detection accuracy and recall when executables are structurally altered or corrupted. Use when the user wants to benchmark on EMBER, or asks about evaluating this task. Reports Recall.

researchpython
0
3
Cesnet Timeseries24 Cl EvalA

This evaluation probes the stability and performance of continual learning algorithms on multivariate time-series forecasting when the underlying data stream is partitioned into tasks with varying temporal granularities. It measures how sensitive forecasting accuracy, catastrophic forgetting, and backward transfer are to the choice of task boundaries and window lengths. Use when the user wants to benchmark on CESNET-Timeseries24, or asks about evaluating this task. Reports Average MSE.

researchpythongo
0
3
Cesped Pose Estimation EvalA

This benchmark evaluates supervised deep learning models for predicting 3D particle orientations (rotation matrices) from 2D Cryo-EM micrographs. It assesses both angular prediction accuracy and the downstream quality of 3D structural reconstructions derived from the predicted poses. Use when the user wants to benchmark on CESPED, or asks about evaluating this task. Reports MAnE.

researchpythonshell
0
3
Ceval EvalA

Evaluates Chinese foundation models' domain knowledge and reasoning capabilities across 52 academic disciplines and four difficulty levels using multiple-choice questions. It probes the models' ability to follow instructions, perform in-context learning, and generate chain-of-thought reasoning in a Chinese language context. Use when the user wants to benchmark on C-EVAL, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cfbenchmark Basic EvalA

Evaluates Chinese large language models on financial text processing capabilities, specifically entity recognition, text classification, and content generation within the financial domain. It tests the models' adaptability using zero-shot and few-shot (3 examples) prompting strategies across eight distinct tasks. Use when the user wants to benchmark on CFBenchmark-Basic, or asks about evaluating this task. Reports F1-Score.

ai-agentspythongo
0
3
Cfbenchmark Mm EvalA

Evaluates multimodal large language models' ability to interpret financial charts, tables, and diagrams in Chinese, and answer domain-specific questions. It probes visual reasoning, statistical and structural analysis, and financial concept comprehension under zero-shot conditions. Use when the user wants to benchmark on CFBenchmark-MM, or asks about evaluating this task. Reports accuracy.

datapythongo
0
3
Cfdb EvalA

This benchmark evaluates machine learning models' ability to detect fraudulent customer activity by analyzing aggregated behavioral patterns and transaction features at the customer level. It probes anomaly detection and risk profiling capabilities on synthetic, privacy-compliant financial data with highly imbalanced class distributions. Use when the user wants to benchmark on CFDB (Customer-level Fraud Detection Benchmark), or asks about evaluating this task. Reports F1 Score.

researchpythongit
0
3
Cfever EvalA

Evaluates a model's ability to retrieve supporting or refuting evidence from Chinese Wikipedia documents and sentences, and subsequently verify claims by predicting their factual status (Supports, Refutes, or Not Enough Info) based on the retrieved evidence. Use when the user wants to benchmark on CFEVER, or asks about evaluating this task. Reports FEVER Score.

researchpythongo
0
3
Cfinbench EvalA

Evaluates large language models' domain-specific knowledge and reasoning in the Chinese financial context, covering foundational knowledge, professional certifications, practical tasks, and regulatory compliance. Use when the user wants to benchmark on CFinBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cfkgr EvalA

Evaluates a model's ability to perform counterfactual reasoning on knowledge graphs by determining whether a target triple remains plausible after a hypothetical scenario is introduced, while also assessing knowledge retention for unaffected facts. Use when the user wants to benchmark on CFKGR-CoDEx-S, CFKGR-CoDEx-M, CFKGR-CoDEx-L, CFKGR-CoDEx-M*, or asks about evaluating this task. Reports Overall F1-score.

researchpythongo
0
3
Cflue EvalA

Evaluates large language models' proficiency in Chinese financial domain knowledge and their ability to perform standard NLP tasks within the financial sector. It probes both factual recall and reasoning via multiple-choice qualification exams, as well as practical application skills like text classification, machine translation, relation extraction, reading comprehension, and text generation. Use when the user wants to benchmark on CFLUE, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cfq EvalA

This benchmark evaluates compositional generalization in semantic parsing by measuring how well models translate anonymized natural language questions into executable SPARQL queries. It specifically probes the ability to generalize to unseen combinations of logical rules (compounds) while maintaining familiarity with individual rules (atoms), using Maximum Compound Divergence splits to ensure fair yet challenging evaluation. Use when the user wants to benchmark on CFQ, or asks about evaluatin...

researchpythongo
0
3
Cfr Retrieval EvalA

Evaluates an information retrieval system's ability to navigate complex, hierarchical, and temporally-varying regulatory documents. It probes the model's capacity to resolve dense cross-references and versioning conflicts to provide complete and accurate answers. Use when the user wants to benchmark on Code of Federal Regulations (CFR), or asks about evaluating this task. Reports Accuracy (Correct/Complete Answers).

researchpythongo
0
3
Cfsl Benchmark EvalA

Evaluates a model's ability to learn sequentially from small, task-specific data batches (continual few-shot learning) without access to prior tasks, measuring sample efficiency and susceptibility to catastrophic forgetting across sequential 5-way 1-shot classification tasks. Use when the user wants to benchmark on Omniglot, SlimImageNet64, or asks about evaluating this task. Reports accuracy.

researchpythontesting
0
3
Cfsl Instance EvalA

Evaluates a model's ability to perform continual few-shot learning and recognize specific object instances under varying class counts and corruption levels. It probes instance-level memorization and robustness to noise and occlusion in a streaming episodic setting. Use when the user wants to benchmark on CFSL synthetic images (SlimageNet64), or asks about evaluating this task. Reports accuracy.

researchpythongit
0
3
Cg Bench EvalA

Evaluates multimodal large language models on clue-grounded audio-visual counting tasks over long videos. It probes the model's ability to integrate audio and visual cues to locate temporal segments and accurately count events, objects, or attributes within those segments. Use when the user wants to benchmark on CG-Bench, or asks about evaluating this task. Reports counting_accuracy.

researchpythongo
0
3
Cgce EvalA

Evaluates Chinese generative chat models on general knowledge and financial domain tasks, measuring response quality across multiple human-assessed dimensions. It probes the model's ability to handle diverse prompts in mathematics, reasoning, scenario writing, and financial analysis, while assessing the overall quality of the generated Chinese text. Use when the user wants to benchmark on CGCE, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Cgiqa 6k EvalA

Evaluates the ability of no-reference image quality assessment (IQA) models to predict human-perceived quality scores for in-the-wild computer graphics images. It probes how well models capture both distortion artifacts and aesthetic quality in synthetic visual content compared to natural scenes. Use when the user wants to benchmark on CGIQA-6k, CCT-CGI, NBU-CIQAD, LIVE-YT-Gaming, or asks about evaluating this task. Reports SRCC.

researchpythongit
0
3
Ch4 Detection Intensity EvalA

Evaluates ensemble machine learning models for binary detection of fugitive methane emissions and continuous prediction of their tracer concentration intensity using meteorological data. Use when the user wants to benchmark on HYSPLIT-generated Methane Emission Dataset, or asks about evaluating this task. Reports accuracy.

researchpythongit
0
3
Chaic EvalA

Evaluates AI agents' ability to socially perceive and cooperatively assist physically constrained humans in long-horizon indoor and outdoor tasks. It probes cooperative planning, goal inference from egocentric visual input, and emergency response under physical constraints. Use when the user wants to benchmark on CHAIC, or asks about evaluating this task. Reports Transport Rate (TR).

researchpythongo
0
3
Chain Of Instructions EvalA

Evaluates LLMs on multi-step compositional instruction following, generalization to hard single-step tasks, and multilingual summarization. It probes the model's ability to chain subtask outputs as inputs for subsequent steps and maintain coherence across multiple instructions. Use when the user wants to benchmark on CoI2-test, CoI3-test, BIG-Bench Hard (BBH), Multilingual Summarization, or asks about evaluating this task. Reports Rouge-L.

researchpythonperformance
0
3
Chainv EvalA

Evaluates the accuracy and inference efficiency of training-free multimodal reasoning methods across diverse vision-language benchmarks. It probes how well atomic visual hint injection reduces redundant reasoning steps while maintaining or improving task performance on math, logic, science, and general visual understanding tasks. Use when the user wants to benchmark on MathVista mini, MathVision, WeMath, MMMU Pro vis, LogicVista, OlympiadBench, VStar, CVBench, ConBench, ChartVQA, SEED-Bench, ...

researchpythongo
0
3
Champkit EvalA

Evaluates the transfer learning capability and generalization of deep learning models (CNNs and ViTs) on patch-level histopathology image classification tasks across multiple cancer-related benchmarks. Use when the user wants to benchmark on Various publicly available histopathology datasets, or asks about evaluating this task. Reports AUROC.

researchpythongit
0
3
Chandassu Metrical EvalA

Evaluates a model's ability to recognize and verify traditional Telugu Chandassu metrical patterns in padyam poetry. It measures adherence to structural prosodic constraints including syllable counts, line divisions, sequential gana patterns, rhythmic breaks, and recurring syllable markers. Use when the user wants to benchmark on Telugu Chandassu Padyam Dataset, or asks about evaluating this task. Reports Chandassu Score.

researchpythongo
0
3
Chanelcolgate Average PrecisionA

Compute chanelcolgate/average_precision via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of chanelcolgate/average_precision.

developmentpython
0
3
Change Detection EvalA

Evaluates a model's ability to detect semantic building changes between two temporally separated remote sensing images. It probes robustness to weak temporal supervision, label noise, and out-of-domain generalization in large-scale urban environments. Use when the user wants to benchmark on b-FLAIR-test, b-FLAIR-test-spot, LEVIR-CD, WHUCD, S2Looking, or asks about evaluating this task. Reports F1-score (F1), Intersection over Union (IoU).

researchpython
0
3
Chaos Forecasting EvalA

Evaluates the ability of time series forecasting models to predict trajectories of low-dimensional chaotic dynamical systems. It probes how well models capture underlying deterministic chaos, smoothness, and multi-scale temporal dependencies without explicit trend or seasonality signals. Use when the user wants to benchmark on Chaotic Dynamical Systems Benchmark, or asks about evaluating this task. Reports sMAPE.

datapythonexpress
0
3
Chaotic Time Series Forecasting EvalA

Evaluates the ability of time series forecasting models to predict future values of noisy, chaotic dynamical systems. It probes how well models capture underlying nonlinear dynamics and handle varying levels of observation noise and system complexity. Use when the user wants to benchmark on Gilpin chaotic systems benchmark, or asks about evaluating this task. Reports SMAPE.

researchpythontesting
0
3
Character Detection MatchingA

Evaluates the visual fidelity and spatial accuracy of formula recognition models by comparing rendered images of predicted and ground-truth LaTeX code at the character level. It addresses the misalignment of text-based metrics with human perception by treating each character as a detectable object in an image. Use when the user has predictions and gold and needs to compute CDM.

researchpythongo
0
3
Character Level F1A

This evaluation probes an LLM-based autorater's ability to predict fine-grained machine translation errors (spans, severities, categories) without using human references. It specifically tests how well the model can specialize to a given test set by leveraging in-context examples of human ratings from other systems on the same inputs. Use when the user has predictions and gold and needs to compute character-level F1.

researchpythongo
0
3
CharerrorrateA

Compute the CharErrorRate metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute CharErrorRate, or asks how to score with CharErrorRate.

documentationpython
0
3
Charged Particle Tracking ResidualsA

Evaluates the accuracy of neural networks in regressing charged particle hit positions (x, y) and incident angles (alpha, beta) from pixelated silicon sensor charge patterns. It measures how well on-chip inference models reconstruct particle trajectories compared to ground-truth simulation and traditional reconstruction algorithms. Use when the user wants to benchmark on Simulated silicon tracker charge clusters, or asks about evaluating this task. Reports residual_68pct_interval.

researchpythongo
0
3
Chart Hqa EvalA

Evaluates multimodal large language models' ability to perform counterfactual reasoning over chart visualizations. It probes whether models rely on parametric memory or truly understand the visual data when answering questions that contain hypothetical assumptions about the chart. Use when the user wants to benchmark on Chart-HQA, or asks about evaluating this task. Reports relaxed accuracy.

researchpythongo
0
3
Chart Qa EvalA

Evaluates vision-language models on chart question answering, probing their ability to extract numerical data, perform visual interpolation, and reason over diverse chart types. It measures both direct data retrieval and complex reasoning capabilities across real-world and synthetic chart distributions. Use when the user wants to benchmark on FigureQA-Sub, DVQA-Sub, PlotQA-Sub, ChartQA, CharXiv, or asks about evaluating this task. Reports exact accuracy, relaxed accuracy.

researchpythongo
0
3
Chart Reasoning EvalA

Evaluates a model's ability to understand complex chart visualizations and perform fine-grained visual grounding and numerical reasoning. It probes both in-domain chart comprehension across real-world and synthetic datasets, and out-of-domain generalization to visual mathematical reasoning tasks. Use when the user wants to benchmark on CharXiv, ChartQAPro, ChartQA, ChartBench, ChartX, ReachQA, MathVista, WeMath, MathVerse, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chart To Code EvalA

Evaluates a model's ability to translate chart images into executable plotting code, measuring both code executability and visual/textual fidelity of the generated charts. Use when the user wants to benchmark on ChartMimic, Plot2Code, ChartX, or asks about evaluating this task. Reports High-Level Score.

researchpythongo
0
3
Chart Understanding EvalA

Evaluates multimodal language models' ability to comprehend diverse chart types, extract underlying numerical data, and answer questions across varying complexity levels. It distinguishes between OCR-dependent recognition on annotated charts and true data reasoning on unannotated or raw-data-requiring charts. Use when the user wants to benchmark on ChartQA, PlotQA, ChartDQA, MMC, ChartX, Chart-to-Table, Chart-to-Text, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chartab EvalA

This benchmark evaluates vision-language models on fine-grained chart understanding, specifically focusing on dense grounding of data and visual attributes, identifying precise differences between paired charts, and measuring robustness to stylistic perturbations like color or font changes. Use when the user wants to benchmark on ChartAB, or asks about evaluating this task. Reports SCRM.

researchpythongo
0
3
Chartassistant EvalA

Evaluates a multimodal language model's ability to comprehend, summarize, and answer questions about various chart types (base and specialized). It probes chart-to-text generation, open-ended and numerical question answering, referring question answering, and chart-to-table translation. Use when the user wants to benchmark on ChartQA, Chart-to-Text, OpenCQA, MathQA, ReferQA, RealQA, or asks about evaluating this task. Reports relaxed_correctness.

researchpythongo
0
3
Chartcap EvalA

Evaluates the capability of vision-language models to generate dense, structurally accurate captions for charts while mitigating hallucinations. It probes both reference-based text similarity and a novel visual consistency metric that verifies if a generated caption can successfully reconstruct the original chart image. Use when the user wants to benchmark on ChartCap, or asks about evaluating this task. Reports Visual Consistency Score.

researchpythongo
0
3
Chartdiff EvalA

This benchmark evaluates a model's ability to perform cross-chart comparative reasoning by generating natural language summaries that identify and explain differences in trends, fluctuations, and anomalies between pairs of charts. It probes vision-language models on their capacity to synthesize visual information from multiple plots into coherent, human-aligned textual descriptions. Use when the user wants to benchmark on ChartDiff, or asks about evaluating this task. Reports GPT Score.

researchpython
0
3
Charte3 EvalA

Evaluates multimodal image-to-image editing models on chart editing tasks, measuring both low-level visual fidelity and high-level semantic correctness and consistency against editing instructions. Use when the user wants to benchmark on ChartE³, or asks about evaluating this task. Reports Correctness.

researchpythongo
0
3
Chartgalaxy EvalA

This benchmark evaluates multimodal large language models' ability to understand and reason about infographic charts. It probes capabilities in text-based data reasoning, visual-element association, and visual style analysis through structured question-answering tasks. Use when the user wants to benchmark on ChartGalaxy, or asks about evaluating this task. Reports relaxed accuracy (5% margin).

researchpythongo
0
3
Chartmimic EvalA

Evaluates large multimodal models' cross-modal reasoning by generating code to reproduce or modify charts based on visual and textual instructions. It tests visual understanding, code generation, and the integration of textual and visual inputs. Use when the user wants to benchmark on ChartMimic, or asks about evaluating this task. Reports Overall.

researchpythongo
0
3
Chartmuseum EvalA

This benchmark evaluates the visual reasoning capabilities of Large Vision-Language Models (LVLMs) on real-world charts. It specifically probes the model's ability to perform visual extraction, object picking, visual comparisons, and trajectory tracking, while distinguishing these from purely textual inference tasks. Use when the user wants to benchmark on CHARTMUSEUM, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Chartnet EvalA

Evaluates multimodal vision-language models on chart understanding tasks, including reconstructing plotting code from charts, extracting tabular data, summarizing chart content, and answering complex reasoning questions. Use when the user wants to benchmark on ChartNet Evaluation Set, or asks about evaluating this task. Reports ChartNet Evaluation Metrics.

researchpythongo
0
3
Chartom EvalA

This benchmark evaluates large language models' ability to comprehend factual data in charts (FACT task) and their capacity to predict how visual manipulations mislead human readers (MIND task). It probes visual theory-of-mind by measuring whether models can distinguish between objective chart data and subjective human perceptual biases introduced by deceptive visualization techniques. Use when the user wants to benchmark on CHARTOM, or asks about evaluating this task. Reports FACT_accuracy.

researchpythongo
0
3