All authors
qhjqhj00 avatar

Claude Skills by qhjqhj00

github.com/qhjqhj00
7,636 skillsA× 7,623B× 11C× 1D× 10 installs2,201 views
Radio Morphology EvalA

Probes the transfer learning capability of self-supervised vision models on radio astronomy morphology classification tasks across heterogeneous imaging pipelines, telescopes, and label granularities. Use when the user wants to benchmark on MiraBest, LoTSS DR2, Radio Galaxy Zoo DR1, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Radio Source Classification EvalA

Evaluates the transferability of self-supervised learning representations for classifying radio astronomy sources from interferometric cutout images. It benchmarks multiple SSL pretraining methods against ImageNet baselines using linear probing and full fine-tuning on downstream classification tasks. Use when the user wants to benchmark on MiraBest, RGZ, MSRS, VLASS, or asks about evaluating this task. Reports classification accuracy.

researchpythongo
0
3
Radioactive EvalA

Evaluates interactive 3D medical image segmentation by measuring how well models segment target structures under varying human-in-the-loop prompting strategies (points, boxes, scribbles) and iterative refinement protocols. It specifically probes the trade-off between interaction effort and segmentation accuracy across 2D and 3D architectures. Use when the user wants to benchmark on RadioActive, or asks about evaluating this task. Reports Dice.

researchpythongit
0
3
Raft Few Shot Classification EvalA

Evaluates few-shot text classification on real-world tasks with class imbalance and long inputs. It probes a model's ability to leverage limited labeled examples, domain knowledge, and open-domain retrieval to classify text without a validation set. Use when the user wants to benchmark on RAFT, or asks about evaluating this task. Reports macro-F1.

researchpythongo
0
3
Raft Optical Flow EvalA

Evaluates dense optical flow estimation accuracy and generalization across synthetic and real-world driving scenes. It measures pixel-wise displacement error and outlier rates on clean and final passes of benchmark datasets. Use when the user wants to benchmark on Sintel, KITTI, or asks about evaluating this task. Reports EPE.

researchpythonperformance
0
3
Rag Coverage EvalA

This evaluation probes the relationship between retrieval effectiveness and downstream information coverage in RAG systems. It measures how well retrieval models capture required information nuggets and how accurately generated responses cover these nuggets with proper citations. Use when the user wants to benchmark on NeuCLIR24, RAG24, WikiVideo, or asks about evaluating this task. Reports Nugget Coverage.

researchpythongo
0
3
Rag Har EvalA

Evaluates a training-free, retrieval-augmented framework for classifying human activities from wearable sensor time-series data. It probes the model's ability to perform open-world activity recognition by retrieving semantically similar sensor examples and using an LLM to predict activity labels without fine-tuning. Use when the user wants to benchmark on HHAR, PAMAP2, MHEALTH, GOTOV, SKODA, USC-HAD, or asks about evaluating this task. Reports accuracy.

ai-agentspythongo
0
3
Rag Medical EvalA

Evaluates the effectiveness and efficiency of Retrieval-Augmented Generation (RAG) systems across medical and general knowledge domains. It probes how different RAG pipeline components (chunking, indexing, query classification, augmentation, and prompting) impact answer accuracy and response latency on question-answering and information extraction tasks. Use when the user wants to benchmark on MMLU, PubMedQA, PromptNER, Query Classification Dataset, or asks about evaluating this task. Reports...

ai-agentspythongo
0
3
Rag Reasoning EvalA

Evaluates retrieval-augmented reasoning systems on their ability to iteratively refine answers using a critique language model. It probes robustness to noisy retrieval, out-of-distribution generalization, and the effectiveness of contrastive critique synthesis over standard self-refinement baselines. Use when the user wants to benchmark on PopQA, TriviaQA, NaturalQuestions, 2WikiMultihopQA, ASQA, HotpotQA, SQuAD, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Rag Robustness EvalA

Evaluates how Retrieval-Augmented Generation (RAG) systems maintain factual accuracy when exposed to adversarial, harmful, or misleading medical evidence. It probes the model's susceptibility to contextual manipulation and its ability to resist misinformation propagation under varying query framings. Use when the user wants to benchmark on TREC Health Misinformation 2020, TREC Health Misinformation 2021, or asks about evaluating this task. Reports ground-truth alignment rate.

researchpythongo
0
3
Rag Tech Docs EvalA

Evaluates how chunking strategies, embedding models, and retrieval thresholds affect RAG performance on technical IEEE documents. Probes the impact of sentence length, keyword position, and acronym handling on retrieval relevance and generator hallucination. Use when the user wants to benchmark on IEEE Wireless LAN MAC/PHY & Battery Glossary, or asks about evaluating this task. Reports qualitative observation.

ai-agentspythonperformance
0
3
Ragen Agent EvalA

Evaluates LLM agents' multi-turn decision-making and reasoning capabilities across symbolic planning, risk-sensitive reasoning, and realistic web interaction environments. It probes the agent's ability to complete interactive tasks under noisy or probabilistic feedback while maintaining exploration and training stability. Use when the user wants to benchmark on Bandit, Sokoban, Frozen Lake, WebShop, or asks about evaluating this task. Reports success rate.

researchpython
0
3
Ragppi EvalA

Evaluates the factual accuracy and atomic fact alignment of LLMs and RAG systems when answering questions about protein-protein interactions (PPIs) in drug discovery. It probes whether models can correctly identify biological, functional, or physical effects between proteins without hallucinating domain-specific details. Use when the user wants to benchmark on RAGPPI, or asks about evaluating this task. Reports F1 (Cosine similarity of atomic facts).

researchpythongit
0
3
Ragsearch EvalA

Evaluates dense RAG and GraphRAG retrieval backends when integrated into agentic search systems. It probes the agent's ability to dynamically retrieve, reason, and answer general and multi-hop QA queries under both training-free prompting and reinforcement learning paradigms. Use when the user wants to benchmark on NQ, PopQA, TriviaQA, HotpotQA, 2Wiki, Musique, or asks about evaluating this task. Reports Exact Match (EM).

ai-agentspythongo
0
3
Raid Robustness EvalA

Evaluates the adversarial robustness and transferability of AI-generated image detectors against crafted perturbations. It probes whether detectors can maintain classification accuracy when faced with white-box and black-box evasion attacks across different perturbation budgets. Use when the user wants to benchmark on RAID, or asks about evaluating this task. Reports F1-score.

researchpythontesting
0
3
Rainbench EvalA

Evaluates deep learning models' ability to forecast global precipitation at multiple lead times (1, 3, 5 days) and estimate same-timestep precipitation using multi-modal satellite and reanalysis data. It probes the model's capacity to handle extreme weather events, class imbalance, and spatial-temporal dependencies in meteorological forecasting. Use when the user wants to benchmark on RainBench, or asks about evaluating this task. Reports Latitude-weighted RMSE.

researchpythongit
0
3
Rainnet EvalA

Evaluates deep learning models for spatial precipitation downscaling by measuring both static reconstruction accuracy and dynamic temporal evolution of rainfall patterns. It probes whether models can capture realistic meteorological properties like heavy rain coverage, cluster movement, and transition speeds. Use when the user wants to benchmark on RainNet, or asks about evaluating this task. Reports PEM.

researchpython
0
3
Rakugo Listening Test EvalA

Evaluates the perceptual quality of synthesized rakugo speech by comparing it to professional human performances across multiple dimensions, including naturalness, character distinguishability, content understandability, entertainment value, and overall skill level. The benchmark probes whether TTS systems can capture the nuanced performance modeling required for traditional Japanese verbal entertainment. Use when the user wants to benchmark on Misomame, or asks about evaluating this task. Re...

researchpythongo
0
3
Ralp Cnn Training EvalA

Evaluates the training throughput and network communication efficiency of distributed CNN training frameworks under varying GPU counts and dataset complexities. Probes how well a system mitigates parameter server bottlenecks and scales across multiple concurrent workloads. Use when the user wants to benchmark on ImageNet-1K, ImageNet-22K, or asks about evaluating this task. Reports throughput (images/sec).

researchpythonnode
0
3
Ram H1200 EvalA

Evaluates medical imaging models on hand radiographs for rheumatoid arthritis. It probes anatomical structure modeling through bone segmentation, fine-grained lesion detection via bone erosion segmentation, and clinical reasoning through ordinal scoring of erosion and joint space narrowing severity. Use when the user wants to benchmark on RAM-H1200, or asks about evaluating this task. Reports DSC, QWK.

researchpythongo
0
3
Rand ScoreA

Compute the rand_score metric — provided by scikit-learn. Use when the user has predictions and ground-truth and needs to compute rand_score, or asks how to score with rand_score.

documentationpython
0
3
RandscoreA

Compute the RandScore metric — provided by torchmetrics. Use when the user has predictions and ground-truth and needs to compute RandScore, or asks how to score with RandScore.

documentationpython
0
3
Randumb Ocl EvalA

Evaluates continual learning methods in online, exemplar-free, and low-exemplar regimes by measuring how well a model retains knowledge of previously seen classes after processing a single pass of sequential data. It specifically tests whether fixed random representations can match or exceed learned representations in these constrained settings. Use when the user wants to benchmark on MNIST, CIFAR10, CIFAR100, TinyImageNet200, miniImageNet100, or asks about evaluating this task. Reports avera...

researchpythonperformance
0
3
Rank Distillm EvalA

Probes the re-ranking capability of distilled cross-encoder models on passage retrieval tasks. It evaluates ranking quality on in-domain benchmarks (TREC Deep Learning tracks) and out-of-domain generalization across diverse corpora (TIREx framework), while also measuring computational efficiency. Use when the user wants to benchmark on Rank-DistiLLM, TREC DL 2019, TREC DL 2020, TIREx, or asks about evaluating this task. Reports nDCG@10.

researchpythongo
0
3
RanksumsA

Compute the ranksums metric — provided by scipy.stats. Use when the user has predictions and ground-truth and needs to compute ranksums, or asks how to score with ranksums.

documentationpython
0
3
Raptorkwok ChinesebleuA

Compute raptorkwok/chinesebleu via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of raptorkwok/chinesebleu.

developmentpython
0
3
Rar B EvalA

Evaluates whether dense retrievers and re-rankers can semantically encode and retrieve correct answers to reasoning problems across diverse tasks. It probes the retriever-LLM behavioral gap by testing performance with and without task instructions, and compares full-dataset retrieval against multiple-choice retrieval settings. Use when the user wants to benchmark on RAR-b, or asks about evaluating this task. Reports nDCG@10.

researchpythongo
0
3
Rasd Medical Image Benchmark EvalA

This evaluation probes the transferability and generalization of medical image foundation models pre-trained exclusively on randomized synthetic data. It measures performance across diverse anatomical regions, imaging modalities (CT, MR, X-ray, ultrasound, fundus), and downstream tasks including segmentation, classification, and detection. Use when the user wants to benchmark on TotalSegmentator, CHAOS, LUNA16, INbreast, STARE, DDTI, or asks about evaluating this task. Reports Dice score, AUC.

researchpythonperformance
0
3
Raspgrade EvalA

Evaluates deep learning models for real-time instance segmentation and ripeness classification of raspberries and punnets in industrial conveyor settings. It probes the model's ability to distinguish between five ripeness grades (OK, Dark, Light, Second, Waste) and background objects under conditions of color similarity and occlusion. Use when the user wants to benchmark on RaspGrade, or asks about evaluating this task. Reports mAP50.

researchpythonperformance
0
3
Ratio Of Stereotypical Responses EvalA

Measures the proportion of times an LLM selects a stereotypical option over anti-stereotype or unrelated alternatives when prompted implicitly or explicitly. Probes the model's susceptibility to implicit bias and its explicit recognition of stereotypes across demographic categories. Use when the user wants to benchmark on StereoSet, CrowSPairs, or asks about evaluating this task. Reports ratio of stereotypical responses.

researchpythongo
0
3
Rationalrewards EvalA

Evaluates a reasoning-based reward model's ability to produce human-aligned preference judgments and optimize visual generation models via reinforcement learning and test-time prompt refinement. Use when the user wants to benchmark on Multimodal Reward Bench 2 (MMRB2), EditReward Bench, GenAI-Bench, ImgEdit-Bench, GEdit-Bench-EN, UniGen (UniGenBench++), PICA-Bench, or asks about evaluating this task. Reports pairwise comparison accuracy.

researchpythongo
0
3
Rats Asr Wer EvalA

This evaluation probes the robustness of automatic speech recognition (ASR) systems when trained on extremely limited in-domain noisy data. It measures how well a model can generalize to real-world noisy conditions by leveraging synthetic noisy data generated via a GAN, compared to traditional data augmentation and fine-tuning baselines. Use when the user wants to benchmark on RATS (Channel A), or asks about evaluating this task. Reports WER (%).

researchpythonexpress
0
3
Rats Channel A Wer EvalA

This evaluation probes the robustness of end-to-end automatic speech recognition systems in noisy acoustic conditions. It measures how well a model preserves speech intelligibility and correctly transcribes utterances when background noise is present, specifically testing the mitigation of over-suppression artifacts during joint speech enhancement and recognition. Use when the user wants to benchmark on RATS Channel-A, or asks about evaluating this task. Reports WER(%).

researchpythonexpress
0
3
Rats Noisy Speech EvalA

Evaluates the fidelity of a simulated noisy speech generator against real VHF/UHF transmitted audio, and measures the downstream robustness of automatic speech recognition (ASR) models trained on the simulated data. Use when the user wants to benchmark on RATS Channel A, or asks about evaluating this task. Reports MSSL, WER.

researchpythontesting
0
3
Ravel EvalA

Evaluates interpretability methods' ability to disentangle polysemantic language model representations by isolating causal attributes through activation interventions on residual stream features. Use when the user wants to benchmark on RAVEL, or asks about evaluating this task. Reports Disentanglescore.

researchpythongit
0
3
Raw Instinct EvalA

Evaluates whether direct classification of RAW sensor data achieves accuracy comparable to traditional RAW-to-RGB converted images, while measuring computational efficiency gains from skipping the conversion pipeline. Use when the user wants to benchmark on Custom RAW/RGB Dataset, or asks about evaluating this task. Reports top-1 classification accuracy.

researchpythonrust
0
3
Rct Numerical Extraction EvalA

Evaluates LLMs' zero-shot capability to classify outcome types and extract precise numerical values from randomized controlled trial reports. It probes the models' numerical reasoning, information extraction robustness, and suitability for automating meta-analysis pipelines. Use when the user wants to benchmark on RCT Numerical Extraction Dataset, or asks about evaluating this task. Reports exact_match_accuracy.

researchpythongo
0
3
Re Laion Caption 19m EvalA

This benchmark evaluates how well text-to-image models adhere to structured prompts by measuring the alignment between generated images and their corresponding captions. It probes the model's ability to preserve semantic details and follow prompt structure during fine-tuning. Use when the user wants to benchmark on Re-LAION-Caption 19M, or asks about evaluating this task. Reports VQA LLaVA.

researchpython
0
3
Re Mi EvalA

Evaluates multimodal LLMs' ability to reason across multiple images, including sequential and set-based consumption, interleaved text-image processing, and heterogeneous visual inputs like charts, equations, maps, and code. It probes cross-image contextual integration, precise visual reading, and step-by-step logical deduction. Use when the user wants to benchmark on ReMI, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Re Verse EvalA

Evaluates vision-language models' ability to comprehend long-form sequential manga narratives, focusing on story synthesis, character grounding, and temporal reasoning across non-linear, multi-panel sequences. Use when the user wants to benchmark on Re:Verse, or asks about evaluating this task. Reports BERTScore.

researchpythongo
0
3
Re2 Peer Review EvalA

Evaluates LLM capabilities across the full academic peer review lifecycle, including predicting paper acceptance and scores, generating structured peer reviews, and simulating multi-turn author-reviewer rebuttal conversations. Use when the user wants to benchmark on Re$^2$, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Reactembed EvalA

Evaluates a cross-domain representation learning framework for protein-molecule interactions by measuring prediction accuracy on regression and classification tasks across diverse biochemical benchmarks. Use when the user wants to benchmark on FreeSolv, CEP, BetaLactamase, Stability, BindingDB, PPIAffinity, BBBP, GO-CC, DrugBank, HumanPPI, YeastPPI, or asks about evaluating this task. Reports Root Mean Square Error (RMSE).

researchpythongo
0
3
Ready Jurist One EvalA

Evaluates the ability of LLM-based agents to perform interactive, procedural legal tasks in dynamic, multi-turn Chinese legal environments. It probes knowledge retrieval, document drafting, and court proceeding navigation, measuring both task completion and adherence to legal procedures. Use when the user wants to benchmark on J1-ENVS, or asks about evaluating this task. Reports average scores.

ai-agentspythongo
0
3
Real 3dqa EvalA

Evaluates whether 3D-LLMs genuinely comprehend 3D spatial relationships rather than relying on linguistic shortcuts or text-only priors. It filters out 3D-independent questions and measures consistency across viewpoint rotations to penalize superficial pattern matching. Use when the user wants to benchmark on Real-3DQA, or asks about evaluating this task. Reports Viewpoint Rotation Score (VRS).

researchpythongo
0
3
Real Iad EvalA

Evaluates industrial anomaly detection models under standard unsupervised and fully unsupervised (noisy training) settings. It probes image-level, pixel-level, and multi-view sample-level defect detection capabilities. Use when the user wants to benchmark on Real-IAD, or asks about evaluating this task. Reports AUROC.

researchpythontesting
0
3
Real Robot Challenge 2022 EvalA

Evaluates offline reinforcement learning and imitation learning algorithms on real-world dexterous manipulation tasks. It probes the ability to learn precise in-hand orientation and stable grasping from pre-collected robot data without online interaction, and measures transfer performance to physical hardware. Use when the user wants to benchmark on Real Robot Challenge 2022 TriFinger Datasets, or asks about evaluating this task. Reports overall score.

researchpythongo
0
3
Real Routing Nco EvalA

Evaluates neural combinatorial optimization models on real-world vehicle routing problems, measuring their ability to generate high-quality routes under asymmetric travel constraints and generalizing to out-of-distribution city maps and location distributions. Use when the user wants to benchmark on Real-World Routing (RRNCO), or asks about evaluating this task. Reports Gap %.

researchpythongit
0
3
Real Time Game Playing EvalA

Evaluates real-time video game control policies across programmatic and real-game environments, measuring task completion, combat effectiveness, human-like behavior, and instruction-following capability. Use when the user wants to benchmark on Hovercraft, Simple-FPS, Real Games (DOOM, Quake, Roblox), or asks about evaluating this task. Reports Hovercraft Loop Time.

researchpythongo
0
3
Real World Cross App EvalA

Assesses multi-step agent capabilities including tool use, GUI grounding, compositional generalization, and long-horizon planning across real-world desktop and web applications. It evaluates whether agents can execute complex, cross-application workflows and self-evaluate their trajectories. Use when the user wants to benchmark on Real-World Cross-Application Benchmark Suite, or asks about evaluating this task. Reports Success.

researchpythonapi
0
3
Realcqa EvalA

Evaluates scientific chart question answering capabilities, specifically testing a model's ability to extract, reason over, and answer questions about real-world scientific charts. It probes first-order logic reasoning and neuro-symbolic capabilities by requiring formal verification of logical inferences from complex visual data. Use when the user wants to benchmark on RealCQA, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3