All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,377
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,601–9,624 of 21,377 skills
- Aste Triplet Extraction EvalEvaluates a model's ability to jointly extract aspect terms, opinion terms, and their sentiment polarities from text. It probes fine-grained aspect-based sentiment analysis by requiring precise span detection and correct pairing of components within sentences. Use when the user wants to benchmark on 14res, 14lap, 15res, 16res, or asks about evaluating this task. Reports F score.Votes: 0GitHub stars: 3
- Ast EvalEvaluates automatic speech translation (AST) and automatic speech recognition (ASR) performance on English-French and English-Romanian datasets. It probes the model's ability to convert spoken source audio directly into written target-language translations, and to recognize speech transcripts. Use when the user wants to benchmark on AST LibriSpeech, MuST-C, or asks about evaluating this task. Reports BLEU.Votes: 0GitHub stars: 3
- Asr4real EvalEvaluates automatic speech recognition (ASR) models on their robustness and fairness across diverse real-world conditions, including accented speech, rehearsed speech, and spontaneous conversational speech. It specifically probes performance disparities related to speaker accent, gender, and socio-economic background. Use when the user wants to benchmark on ALLSSTAR, NISP, VoxPopuli, Buckeye, CORAAL, or asks about evaluating this task. Reports WER.Votes: 0GitHub stars: 3
- Asr Wer EvalEvaluates automatic speech recognition (ASR) performance across English and Croatian by measuring word error rate on multiple held-out test sets. It probes the model's ability to accurately transcribe spoken audio, including handling of punctuation and capitalization. Use when the user wants to benchmark on VoxPopuli, FLEURS, Mozilla Common Voice (MCV12), Hugging Face ASR Leaderboard datasets, or asks about evaluating this task. Reports WER.Votes: 0GitHub stars: 3
- Asr Robustness EvalEvaluates the robustness and cross-domain generalization of Automatic Speech Recognition (ASR) models by measuring Word Error Rate (WER) across multiple public and in-house English speech datasets with varying acoustic conditions, sampling rates, and speech types. Use when the user wants to benchmark on LibriSpeech, SwitchBoard & Fisher, WSJ, Common Voice, TED-LIUM v3, Robust Video, CHiME-6, or asks about evaluating this task. Reports WER.Votes: 0GitHub stars: 3
- Asr Noise Robustness EvalEvaluates the robustness of end-to-end automatic speech recognition models to real-world acoustic distortions, including far-field reverberation, mixed sampling rates, low-bitrate codecs, and background noise at varying signal-to-noise ratios. Use when the user wants to benchmark on LibriSpeech, BUT ReverbDB, Hub5 Switchboard & CallHome, AISHELL-2, or asks about evaluating this task. Reports greedy WER (%).Votes: 0GitHub stars: 3
- Asr Metric Approx EvalEvaluates a label-free regression framework that approximates automatic speech recognition error rates (WER and CER) using multimodal embeddings and predicted transcripts. Probes the tool's robustness across diverse acoustic conditions, domains, and out-of-distribution settings. Use when the user wants to benchmark on LibriSpeech, TED-LIUM, GigaSpeech, SPGISpeech, Common Voice, Earnings22, AMI (IHM), People’s Speech, SLUE-VoXCeleb, Primock57, VoxPopuli Accented, ATCOsim, BERSt, CHiME-6, or as...Votes: 0GitHub stars: 3
- Asr Datasets Benchmarks EvalEvaluates Automatic Speech Recognition systems across diverse acoustic conditions, domains, and linguistic settings. It probes model robustness to read vs. spontaneous speech, clean vs. noisy environments, and demographic bias in multilingual crowdsourced data. Use when the user wants to benchmark on LibriSpeech, Switchboard, TED-LIUM 3, CHiME-6, Common Voice 17.0, or asks about evaluating this task. Reports Word Error Rate (WER).Votes: 0GitHub stars: 3
- Asr Clinical Continual EvalThis evaluation probes an ASR model's ability to continuously adapt to noisy, rural clinical telephony speech while retaining its baseline performance on standard general-domain speech. It specifically measures the trade-off between target-domain transcription accuracy and catastrophic forgetting of pre-trained linguistic knowledge. Use when the user wants to benchmark on Gram Vaani, Kathbath, or asks about evaluating this task. Reports WER.Votes: 0GitHub stars: 3
- Asr Bambara EvalEvaluates automatic speech recognition (ASR) models on spontaneous speech in Bambara, a low-resource West African language. It probes the models' ability to accurately transcribe audio segments in both a controlled test set and a more heterogeneous benchmark. Use when the user wants to benchmark on Afvoices Test, Nyana Eval, or asks about evaluating this task. Reports WER (%).Votes: 0GitHub stars: 3
- Aspire EvalEvaluates the perceptual quality of audio-visual speech enhancement models in real-world noisy environments. It probes how well models generalize to natural reverberation, multi-source background noise, and speaker occlusion compared to synthetic training conditions. Use when the user wants to benchmark on ASPIRE, or asks about evaluating this task. Reports MUSHRA.Votes: 0GitHub stars: 3
- Asped EvalBinary audio classification to detect the presence of pedestrians in urban environments. It probes a model's ability to distinguish pedestrian activity from background noise under varying spatial radii and pedestrian count thresholds. Use when the user wants to benchmark on ASPED, or asks about evaluating this task. Reports macro-average recall.Votes: 0GitHub stars: 3
- Aspect Extraction EvalEvaluates a model's ability to identify aspect terms (targets), their categories, sentiment polarity, and exact character positions within restaurant review sentences in Turkish. Use when the user wants to benchmark on SemEval 2016 Turkish Restaurant Reviews, SemEval 2016 English-Translated Restaurant Reviews, or asks about evaluating this task. Reports F1 score.Votes: 0GitHub stars: 3
- Asnm Tun EvalEvaluates classifier robustness against tunneling and non-payload-based adversarial obfuscations in network traffic. It tests whether models trained on direct attacks can detect obfuscated variants and how training data augmentation with obfuscated samples improves detection. Use when the user wants to benchmark on ASNM-TUN, or asks about evaluating this task. Reports F1-measure.Votes: 0GitHub stars: 3
- Asnm Npbo EvalEvaluates classifier resistance to non-payload-based obfuscation (NPBO) techniques like TCP reordering and retransmissions. It tests detection performance when classifiers are trained without vs. with knowledge of obfuscated attacks. Use when the user wants to benchmark on ASNM-NPBO, or asks about evaluating this task. Reports F1-measure.Votes: 0GitHub stars: 3
- Asnm Cdx 2009 EvalEvaluates the ability of machine learning classifiers to detect network intrusions and adversarial obfuscations using aggregated bidirectional TCP flow features. It probes whether models can distinguish legitimate traffic from direct and obfuscated attacks without relying on packet payloads. Use when the user wants to benchmark on ASNM-CDX-2009, or asks about evaluating this task. Reports F1-measure.Votes: 0GitHub stars: 3
- Askbench EvalEvaluates LLMs' ability to detect intent deficiencies or overconfidence in user queries and request targeted clarification during multi-turn interactive QA. It measures how well models balance asking clarifying questions versus providing final answers, using rubric-based checkpoints to score clarification quality and final answer accuracy. Use when the user wants to benchmark on AskBench, HealthBench, or asks about evaluating this task. Reports single-turn accuracy.Votes: 0GitHub stars: 3
- Asimov Safety EvalThis benchmark evaluates the semantic safety and ethical reasoning of vision-language models in robotics. It probes whether models can correctly identify desirable versus undesirable actions across multimodal scenes, real-world injury scenarios, and hypothetical ethical dilemmas. Use when the user wants to benchmark on ASIMOV, or asks about evaluating this task. Reports classification accuracy.Votes: 0GitHub stars: 3
- Asgardbench EvalThis benchmark evaluates visually grounded interactive planning by testing an agent's ability to dynamically adapt action sequences based on real-time visual observations. It isolates plan adaptation from navigation and low-level manipulation, measuring how well models track environmental state and revise plans under minimal or absent corrective feedback. Use when the user wants to benchmark on AsgardBench, or asks about evaluating this task. Reports success_rate.Votes: 0GitHub stars: 3
- Asd And Se EvalEvaluates a unified audio-visual model's ability to detect which speaker is actively speaking in multi-person video scenes and to enhance speech signals by removing background noise and interference. Use when the user wants to benchmark on AVA-ActiveSpeaker, LRS2, TalkSet, Columbia, MUSAN, or asks about evaluating this task. Reports mAP.Votes: 0GitHub stars: 3
- Asclepius EvalThis benchmark evaluates the clinical reasoning, perception, and diagnostic capabilities of multi-modal large language models across 15 medical specialties. It probes tasks ranging from anatomical and attribute perception to disease identification, staging, treatment planning, and medical report generation. Use when the user wants to benchmark on Asclepius, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Ascend EvalEvaluates automatic speech recognition (ASR) models on spontaneous Mandarin-English code-switching in multi-turn conversations. It probes the model's ability to accurately transcribe mixed-language speech under realistic, unscripted conditions with diverse speaker backgrounds. Use when the user wants to benchmark on ASCEND, or asks about evaluating this task. Reports MER.Votes: 0GitHub stars: 3
- Arxivcap EvalEvaluates large vision-language models' ability to comprehend and generate text for scientific figures. It probes capabilities in single and multi-figure captioning, contextualized captioning using in-context examples, and inferring paper titles from figure-caption sequences. Use when the user wants to benchmark on ArXivCap, or asks about evaluating this task. Reports BLEU-2.Votes: 0GitHub stars: 3
- Artvip EvalEvaluates the visual realism and physical fidelity of articulated digital assets for robot learning. It measures geometric detail, reconstruction quality, visual feature alignment with real-world data, and joint motion accuracy under external forces. Use when the user wants to benchmark on ArtVIP, or asks about evaluating this task. Reports joint displacement discrepancy.Votes: 0GitHub stars: 3