
Claude Skills by qhjqhj00
github.com/qhjqhj00Evaluates session-based recommendation models by predicting the next item in a user's clickstream session. It probes the model's ability to capture latent, non-temporal transition patterns and item order within a session graph. Use when the user wants to benchmark on Yoochoose, Diginetica, or asks about evaluating this task. Reports R@20.
Evaluates a fine-grained selective similarity integration framework for drug-target interaction prediction. It tests the model's ability to dynamically weight multiple drug and target similarity views based on local interaction consistency to predict binary interaction labels. Use when the user wants to benchmark on Nuclear Receptors (NR), G-protein coupled receptors (GPCR), Ion Channel (IC), Enzyme (E), Luo, or asks about evaluating this task. Reports AUC.
Evaluates an LLM verifier's ability to rank multiple candidate answers by their factual correctness and error severity across single-hop and multi-hop questions, using external knowledge retrieval and fine-grained scoring. Use when the user wants to benchmark on FGVeriBench, or asks about evaluating this task. Reports Kendall-tau.
Evaluates the ability of shallow and deep learning models to predict click-through rates (CTR) on large-scale ad impression datasets. It probes how well architectures can model high-order feature interactions and dynamically weight feature importance using bilinear functions and Squeeze-Excitation networks. Use when the user wants to benchmark on Criteo, Avazu, or asks about evaluating this task. Reports AUC.
Measures the distributional similarity between original GAN-generated images and their semantically manipulated counterparts. It evaluates whether a latent space transformation preserves overall image quality and realism while altering specific attributes. Use when the user has predictions and gold and needs to compute FID.
Probes whether input attribution methods accurately reflect token importance for model predictions, particularly when inputs are adversarially perturbed or masked out-of-distribution. It evaluates the consistency of fidelity scores across different model architectures and under various adversarial attacks. Use when the user has predictions and gold and needs to compute fidelity.
Evaluates large language models' ability to follow complex financial instructions, with a strong emphasis on precise adherence to formatting, structural constraints, and conditional styling requirements. It probes whether models can maintain procedural compliance rather than just semantic correctness. Use when the user wants to benchmark on FIFE, or asks about evaluating this task. Reports Strict compliance.
Evaluates language models' ability to interpret nonliteral, creative metaphors by selecting the correct literal meaning from two opposing options, and generating sensible interpretations for novel metaphors. It probes commonsense grounding and contextual understanding beyond literal paraphrase tasks. Use when the user wants to benchmark on Fig-QA, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the ability of vision-language and image-editing models to perform semantically correct, structure-aware modifications to scientific charts. It probes whether models can follow precise editing instructions while preserving data-encoding consistency, axis coherence, and legend integrity, rather than merely producing pixel-level visual similarity. Use when the user wants to benchmark on FigEdit, or asks about evaluating this task. Reports Instruction-following score.
Evaluates the ability of text-to-illustration models to generate publication-ready scientific figures that balance structural fidelity, visual aesthetics, and communicative clarity based on long-form scientific text. Use when the user wants to benchmark on FigureBench, or asks about evaluating this task. Reports Overall score.
Evaluates LLMs' ability to understand and generate text in Filipino, Tagalog, and Cebuano across cultural knowledge, classical NLP tasks, reading comprehension, and text generation. It probes cultural alignment, factual recall, linguistic processing, and translation capabilities in low-resource Southeast Asian languages. Use when the user wants to benchmark on FilBench, or asks about evaluating this task. Reports FilBench Score.
Evaluates AI agents' ability to personalize based on file-system behavioral traces across procedural, semantic, and episodic memory channels. It probes attribute recognition, behavioral inference, anomaly detection, and grounding using simulated, multimodal, and real-world settings. Use when the user wants to benchmark on FileGramBench, or asks about evaluating this task. Reports accuracy.
Evaluates text classification capability in low-resource settings for the Filipino language, specifically measuring model robustness and performance degradation as training data size is systematically reduced. Use when the user wants to benchmark on Hate Speech, Dengue, or asks about evaluating this task. Reports accuracy, hamming loss.
Evaluates a model's ability to detect and classify filler words (e.g., 'uh', 'um') in naturalistic speech recordings. It probes temporal localization accuracy and fine-grained acoustic classification under varying evaluation granularities. Use when the user wants to benchmark on PodcastFillers, or asks about evaluating this task. Reports F1.
Evaluates Finnish large language models on a curated suite of 11 tasks spanning arithmetic, reasoning, general knowledge, emotion classification, and linguistic understanding. It probes few-shot generalization and task-specific capabilities in a low-resource language setting. Use when the user wants to benchmark on FIN-bench, or asks about evaluating this task. Reports mean accuracy.
Evaluates large language models on Finnish language capabilities across reading comprehension, commonsense reasoning, sentiment analysis, world knowledge, truthfulness, and alignment. It probes both multiple-choice and generative capabilities under varying prompt formulations (cloze vs. multiple-choice) and shot configurations. Use when the user wants to benchmark on FIN-bench-v2, or asks about evaluating this task. Reports normalized accuracy.
Evaluates large language models' ability to answer financial questions using various retrieval and context strategies, probing numerical reasoning, factuality, and handling of structured or long documents. Use when the user wants to benchmark on FinanceBench, or asks about evaluating this task. Reports correct answer.
This benchmark evaluates large language models' ability to perform knowledge-intensive mathematical reasoning within the finance domain. It requires models to integrate college-level financial knowledge with both textual descriptions and tabular data to solve complex problems. Use when the user wants to benchmark on FinanceMATH, or asks about evaluating this task. Reports accuracy.
Evaluates the performance, efficiency, and robustness of small language models (≤10B parameters) under three agent paradigms: base prompting, single-agent tool use, and multi-agent collaboration. It measures how architectural complexity impacts accuracy, latency, and reliability across diverse financial tasks. Use when the user wants to benchmark on Financial Agent Benchmark (20 datasets), or asks about evaluating this task. Reports Composite Effectiveness Score.
Evaluates large language models on a suite of 13 financial and ESG-related NLP tasks, including sentiment analysis, classification, named entity recognition, relation extraction, financial question answering, text summarization, and sustainability report generation. It probes the model's ability to perform domain-specific reasoning, information extraction, and structured text generation in the financial sector. Use when the user wants to benchmark on FiQASA, FOMC, MultiFin, MLESG, NER, FINER-...
Evaluates whether LLMs can perform domain-level conceptual understanding and precise financial reasoning across multilingual professional certification exams. It probes analytical rigor and regulatory knowledge integration rather than simple factual recall. Use when the user wants to benchmark on EFPA, GRFinQA, CFA, CPA, BBF, SAHM, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates the zero-shot instruction-following and classification capabilities of large language models on four financial NLP tasks: FOMC communication sentiment, general financial sentiment, numerical claim detection, and named entity recognition. It measures how well generative models can perform these tasks without fine-tuning compared to traditional PLMs. Use when the user wants to benchmark on Financial NLP Tasks (FOMC, Sentiment, Claim Detection, NER), or asks about evalua...
This benchmark evaluates a model's ability to detect financial misinformation in a reference-free setting, where models must classify financial narrative paragraphs as true or false without external knowledge or source documents. It probes semantic pattern recognition for financial manipulation, omission detection, and domain-specific generalization. Use when the user wants to benchmark on MisD@ICWSM2026, or asks about evaluating this task. Reports Accuracy.
Evaluates large language models on ten financial NLP tasks to measure classification and generation accuracy, inference speed, and a novel Token Efficiency Score (TES) that quantifies the performance-compute trade-off. Use when the user wants to benchmark on Financial NLP tasks (10 datasets), or asks about evaluating this task. Reports Token Efficiency Score (TES).
Evaluates the capability of large language models (ChatGPT and GPT-4) and domain-specific models to solve a variety of financial text analytics tasks. It probes performance across sentiment analysis, classification, information extraction, and question answering, measuring how well models handle domain-specific knowledge and structured prediction. Use when the user wants to benchmark on Financial NLP Tasks (Sentiment, Classification, NER, RE, QA), or asks about evaluating this task. Reports a...
This benchmark probes a model's ability to classify financial news sentences or phrases into positive, neutral, or negative semantic orientations. It specifically tests domain-specific sentiment analysis by evaluating how well models capture contextual cues, economic concepts, and directional event expectations in financial texts. Use when the user wants to benchmark on Financial PhraseBank, or asks about evaluating this task. Reports accuracy.
Probes the ability of LLMs and traditional NLP tools to accurately classify financial sentiment (positive, neutral, or negative) from news headlines and earnings-related text. It specifically evaluates how well models capture nuanced, hedged, or domain-specific financial language compared to baseline sentiment engines. Use when the user wants to benchmark on Financial Phrase Bank, or asks about evaluating this task. Reports accuracy.
Evaluates the ability of text embedding models to retrieve relevant financial document passages given complex, long-form queries. It probes domain-specific retrieval capabilities, including sensitivity to company names, tickers, financial metrics, and date-specific information. Use when the user wants to benchmark on Financial Document Retrieval Dataset, or asks about evaluating this task. Reports Recall@1.
Evaluates a model's ability to detect subtle semantic shifts between pairs of financial narratives by ranking similar pairs higher than dissimilar ones. It probes nuanced understanding of financial language, including intensified sentiment, elaborated details, plan realization, and emerging situations. Use when the user wants to benchmark on LLM-augmented FinSTS, Human-annotated FinSTS, or asks about evaluating this task. Reports AUC.
Evaluates multi-step symbolic financial reasoning by measuring how well language models generate verifiable chain-of-thought traces aligned with executable financial templates. It probes both the semantic and numeric consistency of intermediate reasoning steps and the accuracy of the final financial answer. Use when the user wants to benchmark on FinChain, or asks about evaluating this task. Reports ChainEval.
Evaluates vision-language models' ability to comprehend real-world financial charts. It probes spatial reasoning, instruction following, and factual extraction across True/False, Multiple Choice, and open-ended Question Answering tasks. Use when the user wants to benchmark on FinChart-Bench, or asks about evaluating this task. Reports Exact Match (EM).
Evaluates automated interpretability methods on their ability to describe black-box functions across numeric, string, and semantic domains. It probes both static language model capabilities and interactive agent reasoning, including handling complexities like composition, noise, bias, and approximation. Use when the user wants to benchmark on FIND, or asks about evaluating this task. Reports adequately described rate.
Evaluates multi-task and single-task learning models on a curated collection of financial NLP tasks (FinDATA) to assess how task diversity, relatedness, and parameter-efficient architectures impact performance. Use when the user wants to benchmark on FinDATA, or asks about evaluating this task. Reports evaluation metrics in Table 2.
This benchmark evaluates long-term memory and spatio-temporal reasoning in embodied agents. It requires agents to recall specific past interactions from a video history to select goal frames and navigate to target entities in dynamic, photorealistic environments over long-horizon tasks. Use when the user wants to benchmark on FindingDory, or asks about evaluating this task. Reports LL-SR.
Probes the ability of vision-language models to distinguish visually similar subcategories within broader classes (e.g., specific bird species or car models) using a multiple-choice format. Use when the user wants to benchmark on ImageNet-1K, Oxford Flowers-102, Oxford-IIIT Pet-37, Food-101, or asks about evaluating this task. Reports accuracy.
This benchmark evaluates large language models' capacity for fine-grained information extraction under augmented instructions. It specifically probes two generalization capabilities: adapting to unseen information types and handling unfamiliar task forms. Performance is measured using standard extraction metrics to compare encoder-decoder and decoder-only architectures. Use when the user wants to benchmark on Fine-grained Information Extraction Benchmark, or asks about evaluating this task. R...
Evaluates the interpretability and faithfulness of neural NLP models by measuring how well token-level saliency rationales align with human-annotated ground truth rationales and model predictions across sentiment analysis, semantic textual similarity, and machine reading comprehension tasks. Use when the user wants to benchmark on Fine-grained Interpretability Benchmark (SA/STS/MRC), or asks about evaluating this task. Reports Token-F1, MAP.
Evaluates multi-modal large language models on fine-grained visual recognition (FGVR) tasks across six standard datasets. It probes the model's ability to distinguish visually similar sub-categories in both closed-world (seen categories) and open-world (unseen categories) settings using chain-of-thought reasoning. Use when the user wants to benchmark on CaltechUCSD Bird-200, Stanford Car-196, Stanford Dog-120, Flower-102, Oxford-IIIT Pet-37, FGVC-Aircraft, or asks about evaluating this task. ...
Probes the effectiveness of a large-scale text-to-image fine-tuning dataset in improving generation quality, text-image alignment, and instruction following across different model architectures (diffusion and autoregressive). Use when the user wants to benchmark on Artificial Analysis Image Arena (Eval Subset), or asks about evaluating this task. Reports human_win_rate.
Evaluates a model's ability to perform fine-grained fact verification on dialogue responses by checking individual atomic facts against external knowledge. It probes whether models can correctly classify facts as supporting, refuting, or lacking sufficient information, particularly in the presence of hallucinations and imbalanced label distributions. Use when the user wants to benchmark on FineDialFact, or asks about evaluating this task. Reports F1-score.
Evaluates a model's ability to perform sentence-level fact verification on generated summaries and localize specific factuality error types. It measures how well the model's judgments align with human annotations across sentence, summary, and system levels. Use when the user wants to benchmark on FineSumFact, or asks about evaluating this task. Reports balanced accuracy (bAcc).
Evaluates the downstream quality of multilingual pretraining corpora by training language models on them and measuring performance on a standardized suite of fine-tuning tasks across Arabic, Hindi, and Turkish. Use when the user wants to benchmark on FineTasks, or asks about evaluating this task. Reports FineTasks scores.
Evaluates large language models' knowledge and reasoning capabilities in the Chinese financial domain across multiple academic subjects like Finance, Economy, Accounting, and professional Certificates. It tests performance under zero-shot, few-shot, answer-only, and chain-of-thought prompting settings. Use when the user wants to benchmark on FinEval, or asks about evaluating this task. Reports accuracy.
Evaluates vision-language models on a diverse suite of 11 multimodal benchmarks covering visual question answering, chart understanding, document parsing, and general multimodal reasoning. Additionally probes GUI/agentic capabilities on screen interaction tasks. Use when the user wants to benchmark on AI2D, ChartQA, DocVQA, InfoVQA, MME, MMMU, ScienceQA, MMStar, OCRBench, TextVQA, SEED-Bench, Screenspot-V2, Screenspot-Pro, or asks about evaluating this task. Reports mean normalized performanc...
Evaluates the impact of different pre-training data processing steps on downstream model quality by training small language models and measuring performance on a curated suite of multilingual zero-shot benchmarks. Use when the user wants to benchmark on FineWeb2 Early-Signal Benchmark Suite, or asks about evaluating this task. Reports per-category macro-average score.
Evaluates LLMs on tabular financial fraud detection by testing their ability to classify transactions using retrieval-augmented in-context learning and importance-guided feature reduction. It probes whether providing a compact set of high-impact features and relevant historical examples improves classification under severe class imbalance. Use when the user wants to benchmark on CCF, CCFRAUD, IEEE-CIS, PAYSIM, or asks about evaluating this task. Reports F1-score, Matthews Correlation Coeffici...
Evaluates forward-looking argument generation in finance across text-to-claim, chart-to-argument, and news-to-argument tasks. Probes a model's ability to generate plausible, structured future scenarios and claims based on financial inputs while maintaining factual consistency and handling financial terminology and numerals. Use when the user wants to benchmark on FinGen, or asks about evaluating this task. Reports ROUGE-1.
Evaluates the robustness and persistence of natural language fingerprints embedded in LLMs against downstream fine-tuning, quantization, pruning, and model merging. It also measures the initial effectiveness of fingerprint elicitation and the harmlessness to baseline model capabilities. Use when the user wants to benchmark on Alpaca-GPT4, ShareGPT, Dolly 2, or asks about evaluating this task. Reports FSR.
Evaluates a financial domain-specific LLM (FinGPT) across six core NLP tasks: sentiment analysis, text classification, named entity recognition, financial question answering, stock movement prediction, and text summarization. It probes the model's ability to handle domain-specific terminology, numerical reasoning, and structured output generation under instruction-tuning. Use when the user wants to benchmark on FLARE-FPB, FLARE-FIQASA, FinGPT Headline Classification, FinGPT/fingpt-ner, ConvFi...
Evaluates AI models' ability to assess the quality of financial information disclosure in Chinese. It probes four capabilities: identifying relevant questions, determining question relevance, evaluating answer readability, and measuring answer relevance to the question. Use when the user wants to benchmark on FinTruthQA, or asks about evaluating this task. Reports Accuracy, Precision, Recall, F1-score, Micro/Macro/Weighted F1, Quadratic Weighted Kappa (QWK).