
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates the financial natural language processing capabilities of encoder-only and decoder-only language models across multiple classification tasks. It probes zero-shot prompting, in-context learning strategies, and the impact of data availability (public vs. proprietary) on model performance. Use when the user wants to benchmark on FinSent, FPB, FiQA SA, ESG, FLS, QA, Headlines-PDU, Headlines-PDC, Headlines-PDD, Headlines-PI, Headlines-AC, Headlines-FI, Headlines-PS, NER, FOMC, or asks ab...
Evaluates financial multi-modal reasoning capabilities of models on chart-based analysis and domain-specific knowledge. It probes perception, analysis, and reasoning across 18 financial domains and 6 asset classes using multiple-choice and computational problems. Use when the user wants to benchmark on FinMME, or asks about evaluating this task. Reports FinScore.
Evaluates the multimodal financial reasoning capabilities of LLMs and MLLMs on expert-level question-answer pairs spanning 15 financial domains. It probes the models' ability to interpret complex visual data (charts, tables), apply domain-specific formulas, and perform multi-step logical calculations. Use when the user wants to benchmark on FinMR, or asks about evaluating this task. Reports accuracy.
Evaluates embedding models on a finance-specific benchmark across seven standard embedding tasks to measure performance relative to general-domain benchmarks. It probes whether general-purpose models suffer a significant performance drop on domain-specific financial text and investigates if this gap is driven by domain shift or inherent dataset complexity. Use when the user wants to benchmark on FinMTEB, or asks about evaluating this task. Reports FinMTEB Score.
Evaluates the inference performance of binarized neural networks (BNNs) accelerated on FPGAs using the FINN framework. It measures classification throughput, latency, and accuracy across standard image datasets to assess hardware efficiency and resource utilization. Use when the user wants to benchmark on MNIST, CIFAR-10, SVHN, or asks about evaluating this task. Reports classification throughput (FPS).
Evaluates the performance, power, and resource efficiency of quantized neural networks deployed on various FPGA platforms using the FINN-R framework. It probes the trade-offs between network precision, hardware resource usage, throughput, and classification accuracy across embedded and datacenter-scale hardware. Use when the user wants to benchmark on MNIST, CIFAR-10, GTSRB, SVHN, VOC 2007, ImageNet, or asks about evaluating this task. Reports Top-1 Accuracy.
Evaluates named entity recognition systems on Finnish text, testing their ability to identify and classify entities (person, location, organization, product, event, date) in both in-domain news and out-of-domain Wikipedia corpora. It specifically probes domain generalization and the handling of nested entity spans. Use when the user wants to benchmark on Finnish News Corpus, or asks about evaluating this task. Reports F1-score.
Evaluates Finnish NLP capabilities across four core tasks: part-of-speech tagging, named entity recognition, dependency parsing, and text classification on domain-specific and out-of-domain corpora. It probes a model's ability to handle morphologically rich language and varying text registers. Use when the user wants to benchmark on Finnish NLP Benchmarks, or asks about evaluating this task. Reports Average.
Evaluates computer vision models on forest scene understanding, specifically testing instance segmentation, panoptic segmentation, and depth completion in unstructured, densely populated natural environments. Use when the user wants to benchmark on FinnWoodlands, or asks about evaluating this task. Reports mAP@50.
Evaluates the ability of foundation models and tabular baselines to predict financial risk (default, fraud, churn) by classifying customer profiles generated from tabular data. It probes how well profile-based tuning captures richer customer semantics compared to isolated table-based classification. Use when the user wants to benchmark on FinBench, or asks about evaluating this task. Reports F1-score.
Evaluates the quality, schema compliance, diversity, and factual grounding of automatically extracted financial knowledge graph triples from SEC 10-K filings. It measures rule-based compliance, entity and relation coverage, semantic diversity via entropy, and comparative quality using an LLM-as-a-Judge framework. Use when the user wants to benchmark on S&P 100 SEC 10-K Filings (2024), or asks about evaluating this task. Reports CheckRules.
Evaluates financial LLMs across seven text-based tasks (sentiment analysis, NER, number understanding, summarization, stock movement prediction, credit scoring, firm disclosure) and three multimodal/hallucination tasks (ChartQA, FinVQA, FinTerms). It measures domain-specific reasoning, instruction following, and hallucination mitigation in financial contexts. Use when the user wants to benchmark on FinSet, ChartQA, FinVQA, FinTerms-MCQ, FinTerms-Gen, Finance Bench, or asks about evaluating th...
Evaluates the quality of a machine-translated extractive QA dataset (FinSQuAD) by training and testing QA models on it, comparing performance against other translated SQuAD datasets and the original English version. It also assesses translation fidelity through backtranslation and manual error analysis. Use when the user wants to benchmark on Finnish SQuAD2.0, SQuAD2.0, or asks about evaluating this task. Reports exact match (EM).
Evaluates large language models on structure-aware XBRL tagging for financial information. It probes two subtasks: numeric entity identification (FinNI) and fine-grained concept linking (FinCL) against the US-GAAP taxonomy, testing the model's ability to extract structured facts and align them with hierarchical financial concepts. Use when the user wants to benchmark on FinTagging, or asks about evaluating this task. Reports macro-F1.
Evaluates a multi-agent financial system's ability to answer real-world investor inquiries across macroeconomic, industry, and company analysis scenarios. Probes accuracy, thoroughness, clarity, and professional financial reasoning through automated LLM-judging and human preference testing. Use when the user wants to benchmark on NGA Grand Era Investor Inquiries, or asks about evaluating this task. Reports Overall Score.
Evaluates the ability of Large Vision-Language Models (LVLMs) to generate accurate, fluent, and comprehensive captions for long videos. It probes semantic alignment, event coverage, and hallucination resistance by comparing model outputs against multi-annotator human references. Use when the user wants to benchmark on FIOVA, or asks about evaluating this task. Reports FIOVA-DQ F1.
Evaluates autonomous coding agents on their ability to rediscover established scientific findings by autonomously planning, implementing, and executing experiments from scratch based only on high-level research questions. It probes end-to-end research workflow capabilities, including experimental design, code generation, and evidence-based conclusion formation. Use when the user wants to benchmark on FIRE-Bench, or asks about evaluating this task. Reports F1.
Evaluates multimodal geospatial models on wildfire risk prediction across in-distribution and out-of-distribution regions. It probes the ability of vision-language models to generate chain-of-thought reasoning traces that condition a vision decoder for accurate, interpretable spatial risk raster generation. Use when the user wants to benchmark on FireScope-Bench, or asks about evaluating this task. Reports ROC AUC, QWK.
Evaluates the performance and scalability of a federated inference scheduling framework under varying request loads, measuring throughput, latency, and auto-scaling capabilities on distributed HPC resources. Use when the user wants to benchmark on ShareGPT, or asks about evaluating this task. Reports Request throughput (req/s).
Evaluates the ranking effectiveness and inference latency of single-token decoding for listwise document reranking. It probes whether using only the first-token logits of alphabetical identifiers can accurately rank candidate documents compared to full sequence generation or traditional language modeling objectives. Use when the user wants to benchmark on TREC DL19-22, BEIR, MS MARCO, or asks about evaluating this task. Reports ranking effectiveness.
Evaluates speech synthesis models on intelligibility, speaker similarity, and long-form generation across multiple languages. It also assesses subjective qualities like naturalness, instruction-following, and human-level indistinguishability using automated LLM-as-a-Judge and Audio Turing Test frameworks. Use when the user wants to benchmark on Seed-TTS-Eval, CV3-Eval, Minimax Multilingual Testset, Long-TTS-Eval, Audio Turing Test, Emergent TTS Eval, or asks about evaluating this task. Report...
Evaluates computer vision models on fine-grained species classification, multi-label trait identification, and pixel-level trait segmentation in fish images. Probes capabilities in handling long-tailed distributions, out-of-distribution generalization to unseen species, and localizing small/rare anatomical features. Use when the user wants to benchmark on Fish-Vista, or asks about evaluating this task. Reports macro-averaged F1-score, Mean Average Precision (mAP), mean Intersection over Union...
Compute the fisher_exact metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute fisher_exact, or asks how to score with fisher_exact.
This benchmark probes a model's ability to perform pixel-wise anomaly detection and uncertainty estimation in complex urban driving scenes. It specifically measures how well a segmentation wrapper identifies out-of-distribution objects (e.g., lost & found items, static blends, web overlays) without degrading the underlying semantic segmentation accuracy. Use when the user wants to benchmark on Fishyscapes benchmark, or asks about evaluating this task. Reports AP.
Evaluates federated learning methods for medical image segmentation under non-IID data distributions, measuring segmentation accuracy and robustness across multiple clinical tasks and imaging modalities. It compares generic and personalized FL approaches against local training baselines to assess client drift, fairness, and generalization. Use when the user wants to benchmark on Fed-Vessel, Fed-Prostate, Fed-COSAS, Fed-BUS, Fed-MG, Fed-Polyp, Fed-Pancreas, Fed-M&Ms, FeTS2022, or asks about ev...
Evaluates large reasoning models on automatically verifiable textual problem-solving tasks, including academic coursework, word puzzles, cipher deciphering, and algorithmic coding. It probes the models' ability to follow instructions, perform logical deduction, and produce correctly formatted final answers under varying reasoning effort settings. Use when the user wants to benchmark on FlagEval Textual, or asks about evaluating this task. Reports accuracy.
Evaluates federated learning models on real-world, non-IID image data with user-level heterogeneity and long-tailed label distributions. It probes how model convergence and multi-label classification performance degrade under privacy constraints (differential privacy) and distributed training compared to centralized baselines. Use when the user wants to benchmark on FLAIR, or asks about evaluating this task. Reports averaged precision (AP).
Evaluates the ability of models to perform high-resolution land-cover semantic segmentation on aerial imagery. It probes robustness to spatial, temporal, and multi-sensor domain shifts, as well as handling radiometric inconsistencies and phenological variations across diverse landscapes. Use when the user wants to benchmark on FLAIR-one, or asks about evaluating this task. Reports mIoU.
Evaluates semantic segmentation models for fine-grained land cover classification and crop type mapping using multi-sensor remote sensing imagery. It probes the model's ability to fuse spatial, spectral, and temporal modalities (aerial RGBI, SPOT, Sentinel-1/2, DEM) for pixel-level prediction at 20 cm resolution. Use when the user wants to benchmark on FLAIR-HUB, or asks about evaluating this task. Reports mIoU.
Evaluates financial domain knowledge and certification exam readiness in Chinese and English. It probes models' ability to answer multiple-choice questions across 14 professional financial certifications with varying difficulty levels. Use when the user wants to benchmark on FLAME-Cer, or asks about evaluating this task. Reports accuracy rate.
Evaluates foundation and reasoning-reinforced language models across 20 core financial NLP tasks. It probes capabilities in numeric reasoning, entity classification, information retrieval, question answering, and summarization, while also measuring inference efficiency and cost. Use when the user wants to benchmark on FLaME, or asks about evaluating this task. Reports F1.
Evaluates federated learning algorithms for decentralized robotic manipulation across heterogeneous environments. It probes a model's ability to generalize from distributed, non-IID demonstrations under visual and physical perturbations, measuring both action prediction fidelity and task completion success. Use when the user wants to benchmark on FLAME, or asks about evaluating this task. Reports RMSE.
Assesses practical financial application capabilities across a hierarchical framework of 10 primary and 21 secondary scenarios. It probes models' ability to perform real-world financial tasks such as compliance checking, document generation, risk control, data extraction, and client analysis. Use when the user wants to benchmark on FLAME-Sce, or asks about evaluating this task. Reports usability rate.
Evaluates the instruction-tuning effectiveness of models trained on the Flan 2022 collection across held-in, chain-of-thought, and held-out benchmarks. It probes zero-shot and few-shot generalization capabilities on reasoning, knowledge, and natural language understanding tasks. Use when the user wants to benchmark on MMLU, BBH, or asks about evaluating this task. Reports zero-shot/few-shot accuracy.
Evaluates bilingual (Spanish-English) financial understanding, prediction, and generation capabilities of LLMs. Probes cross-lingual transfer, domain-specific instruction following, and performance disparity between high-resource and low-resource financial tasks. Use when the user wants to benchmark on FLARE-ES, or asks about evaluating this task. Reports Acc, F1.
Evaluates the effectiveness of a frequency-domain-guided KV cache compression method (FlashCache) on multimodal long-context understanding tasks. It measures how well the model preserves accuracy under varying KV cache retention ratios and quantifies the computational overhead and decoding latency compared to baseline eviction methods. Use when the user wants to benchmark on MileBench, MUIRBench, MMMU, V*, HR-Bench, FAVOR-Bench, or asks about evaluating this task. Reports accuracy.
Evaluates the effectiveness of various Retrieval-Augmented Generation (RAG) methods across text and multimodal question-answering tasks. It probes how different retrieval strategies, context compression techniques, and generator optimizations impact answer accuracy and faithfulness on single-hop and multi-hop datasets. Use when the user wants to benchmark on NQ, TriviaQA, HotpotQA, 2WikiMultihopQA, Gaokao-MM, MultimodalQA, MathVista, or asks about evaluating this task. Reports Acc.
Evaluates the robustness and efficiency of text-guided visual token pruning in large multimodal models. It probes whether aggressive token compression (retaining 32–128 tokens for images, 114–455 for video) degrades performance on image and video question-answering tasks, and measures cross-modal grounding quality via spatial alignment and semantic overlap metrics. Use when the user wants to benchmark on VQAv2, GQA, VizWiz, ScienceQA-IMG, TextVQA, POPE, MME, MMBench, MMBench-CN, MM Vet, TGIF-...
Evaluates LLMs on fine-grained alignment capabilities by decomposing instruction-following performance into 12 sub-skills across four domains (Logical Thinking, Background Knowledge, Problem Handling, User Alignment). It measures how well models adhere to specific quality criteria like factuality, logical robustness, and harmlessness on a per-instance basis. Use when the user wants to benchmark on FLASK, or asks about evaluating this task. Reports FLASK skill score.
Evaluates multi-agent coordination and path planning for train rescheduling in a dynamic grid-world railway simulation. It probes the ability of agents to adapt to partial observability, handle congestion, and coordinate under switching constraints and dynamic disruptions. Use when the user wants to benchmark on Flatland Competition 2020, or asks about evaluating this task. Reports overall score.
Evaluates the ability of quantum and classical models to perform full-waveform inversion (FWI) by predicting subsurface velocity maps from scaled seismic waveform data. It probes the effectiveness of physics-guided data scaling and layer-wise variational quantum circuit designs in geophysical imaging tasks. Use when the user wants to benchmark on FlatVelA, or asks about evaluating this task. Reports SSIM.
Evaluates a unified vision-language foundation model across 35 downstream tasks spanning vision classification, natural language understanding, and multimodal reasoning/retrieval. It probes the model's ability to generalize from joint unimodal and multimodal pretraining to zero-shot and fine-tuned downstream settings. Use when the user wants to benchmark on GLUE (MNLI, CoLA, MRPC, QQP, SST-2, QNLI, RTE, STS-B), 22 Vision Datasets (ImageNet, Food101, CIFAR10, CIFAR100, Cars, Aircraft, DTD, Pet...
Compute the FleissKappa metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute FleissKappa, or asks how to score with FleissKappa.
Evaluates a multimodal LLM's ability to perform visual reasoning across heterogeneous medical modalities (2D images, 3D volumes, videos) and generate clinical reports. It probes diagnostic accuracy, cross-modal generalization, temporal understanding, and structured medical knowledge integration. Use when the user wants to benchmark on OmniMedVQA, PMC-VQA, VQA-RAD, PathVQA, SLAKE, MIMIC-CXR, IU-Xray, M3D-VQA, MedVideoBench, or asks about evaluating this task. Reports accuracy, ROUGE-L, CIDEr.
This evaluation probes a model's ability to perform automatic speech recognition in low-resource and zero-supervised settings by leveraging joint speech-text representation learning. It specifically measures how well the model can transcribe unseen languages using only untranscribed audio and graphemic text, without relying on manually labeled speech data. Use when the user wants to benchmark on FLEURS, or asks about evaluating this task. Reports CER.
Evaluates universal speech representations across 102 languages using few-shot learning on parallel speech data. Probes capabilities in automatic speech recognition (ASR), speech language identification, and retrieval tasks. Use when the user wants to benchmark on FLEURS, or asks about evaluating this task. Reports character level error rate.
Evaluates multilingual spoken language understanding (SLU) across 102 languages for topical classification and 92 languages for spoken multiple-choice QA, testing cross-lingual transfer, speech-to-text translation, and robustness to audio quality variations. Use when the user wants to benchmark on SIB-Fleurs, Belebele-Fleurs, or asks about evaluating this task. Reports accuracy.
This evaluation probes the inference throughput and generation quality of LLMs across diverse hardware and software configurations. It also tests a predictive modeling framework designed to optimize system co-design by forecasting performance metrics based on model and hardware features. Use when the user wants to benchmark on OpenOrca, Open MLPerf Dataset, or asks about evaluating this task. Reports Tokens/s.
Evaluates the ability of large protein language models to predict protein fitness under constrained, low-data scenarios. It probes mutation-level generalization, overfitting risks, and the impact of model depth and structural information on predictive accuracy across diverse protein families. Use when the user wants to benchmark on FLIP benchmark, or asks about evaluating this task. Reports MSE.
Evaluates a novel tiling and augmentation strategy for Earth observation imagery against conventional tiling. It measures the method's ability to preserve spatial context and improve semantic segmentation performance on highly imbalanced geospatial data. Use when the user wants to benchmark on Land Cover of Canada (LCC), or asks about evaluating this task. Reports precision.