All categories
Research
Research, evidence gathering, literature, reports, investigation, and synthesis
- 21,377
- 891
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browserBrowse research skills
Showing 9,505–9,528 of 21,377 skills
- Babelcode EvalEvaluates the capability of large language models to generate executable code and translate code across multiple programming languages. It measures functional correctness by executing generated programs against test cases and computing the probability that at least one sample passes. Use when the user wants to benchmark on BC-HumanEval, BC-MBPP, BC-Transcoder, TP3, or asks about evaluating this task. Reports pass@k.Votes: 0GitHub stars: 3
- B2t2 EvalEvaluates the expressiveness and type-error diagnostic quality of tabular programming type systems. It tests whether a type system can correctly type a curated set of table operations, handle example programs, and provide accurate feedback on buggy code. Use when the user wants to benchmark on Example Tables, or asks about evaluating this task. Reports expressiveness.Votes: 0GitHub stars: 3
- Ayah Alignment Coverage EvalEvaluates the ability of an audio segmentation pipeline to correctly identify and align individual Quranic verses (ayahs) from long-form recitations. It probes the robustness of alignment methods and ASR backbones against recitation style variations and phonological differences. Use when the user wants to benchmark on Tadabur Evaluation Set (5 Reciters), or asks about evaluating this task. Reports Alignment Coverage (%).Votes: 0GitHub stars: 3
- Aya EvalEvaluates the open-ended generation capabilities of multilingual LLMs, including brainstorming, planning, and unstructured long-form responses across diverse languages, scripts, and resource levels. Use when the user wants to benchmark on AYA-HUMAN-ANNOTATED, DOLLY-MACHINE-TRANSLATED, DOLLY-HUMAN-EDITED, or asks about evaluating this task. Reports fluency and quality.Votes: 0GitHub stars: 3
- Axonn Scaling EvalEvaluates the scaling efficiency and hardware utilization of asynchronous deep learning frameworks (AxoNN, Megatron-LM, DeepSpeed) on large-scale transformer models. It measures how well frameworks overlap communication and computation across varying GPU counts and model sizes while training on a fixed text corpus. Use when the user wants to benchmark on wikitext-103, or asks about evaluating this task. Reports expected_training_time.Votes: 0GitHub stars: 3
- Avut EvalEvaluates multimodal large language models on their ability to comprehend audio content within videos and align audio cues with corresponding visual information. It specifically probes whether models rely on genuine multimodal reasoning or fall back to text-based shortcuts. Use when the user wants to benchmark on AVUT, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Avse Cog Mhear EvalEvaluates audio-visual speech enhancement models on their ability to suppress background noise and competing speakers while preserving speech intelligibility and perceptual quality in real-time hearing aid scenarios. Use when the user wants to benchmark on COG-MHEAR AVSE Challenge, or asks about evaluating this task. Reports PESQ.Votes: 0GitHub stars: 3
- Avsd Dst EvalEvaluates a model's ability to perform dialogue state tracking in an open-domain, multimodal setting by framing it as a question-answering task. It measures how well the system tracks conversation state and generates accurate answers based on video/audio context and dialogue history. Use when the user wants to benchmark on AVSD, or asks about evaluating this task. Reports F1 score.Votes: 0GitHub stars: 3
- Avsbench Vpo EvalThis evaluation probes an audio-visual segmentation model's ability to accurately localize and segment visual objects that correspond to sounding audio sources. It specifically tests robustness across single-source, multi-source, and semantically ambiguous scenarios where visual distractors or overlapping sounds may be present. Use when the user wants to benchmark on AVSBench, VPO, or asks about evaluating this task. Reports J&F.Votes: 0GitHub stars: 3
- Avrobustbench EvalEvaluates the robustness of audio-visual recognition models when subjected to simultaneous, correlated corruptions across both audio and video modalities at test-time. It measures how well models maintain classification accuracy under 75 distinct bimodal distributional shifts ranging from mild to extreme severity. Use when the user wants to benchmark on AudioSet-2C, VGGSound-2C, Kinetics-2C, EpicKitchens-2C, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Avid EvalEvaluates a model's ability to detect, classify, and temporally ground audio-visual inconsistencies in long-form videos, as well as generate causal explanations for cross-modal mismatches across eight fine-grained categories. Use when the user wants to benchmark on AVID, or asks about evaluating this task. Reports mIoU.Votes: 0GitHub stars: 3
- Avgen Bench EvalEvaluates the multi-granular capabilities of Text-to-Audio-Video (T2AV) generation models across basic uni-modal fidelity, cross-modal synchronization, and fine-grained dimensions. It probes specific capabilities including text rendering, facial consistency, musical pitch control, speech coherence, and physical plausibility to reveal systematic failure modes in current generative systems. Use when the user wants to benchmark on AVGen-Bench, or asks about evaluating this task. Reports Total.Votes: 0GitHub stars: 3
- Averitec EvalEvaluates a model's ability to verify real-world claims by retrieving web evidence, generating supporting questions, predicting veracity stance, and producing textual justifications. It probes retrieval quality, stance detection, and justification generation under realistic conditions with temporal and context constraints. Use when the user wants to benchmark on AVeriTeC, or asks about evaluating this task. Reports Macro-F1.Votes: 0GitHub stars: 3
- Averimavec EvalThis benchmark probes a model's ability to perform multimodal fact-checking by verifying real-world image-text claims. It requires the system to retrieve cross-modal evidence, analyze inconsistencies, and produce a justified verdict that aligns with ground truth labels. Use when the user wants to benchmark on AVerImaTeC, or asks about evaluating this task. Reports verdict_correctness.Votes: 0GitHub stars: 3
- Avere Emotion Reasoning EvalThis benchmark probes multimodal large language models' ability to reason about emotions from audio and video inputs while avoiding spurious cue associations and hallucinations. It specifically tests whether models can correctly align relevant audiovisual cues with emotional labels and resist over-reliance on textual priors or irrelevant modalities. Use when the user wants to benchmark on EmoReAlM, DFEW, RAVDESS, MER2023, EMER, or asks about evaluating this task. Reports average accuracy.Votes: 0GitHub stars: 3
- Average Per Token Log ProbEvaluates language models on multiple-choice or candidate-selection downstream tasks by scoring candidate answers based on their likelihood under the model. It measures how well the model assigns high probability to the correct answer among a set of options. Use when the user has predictions and gold and needs to compute average_per_token_log_prob.Votes: 0GitHub stars: 3
- Average Frame TimingThis benchmark evaluates the real-time rendering performance of a VR NeRF system by measuring the time required to generate each frame under varying field-of-view (FoV) and pixel-per-degree (PPD) settings. It probes the system's ability to maintain interactive framerates (≥30 FPS) while fusing neural radiance fields with CAD geometry in immersive virtual reality. Use when the user has predictions and gold and needs to compute average frame timing.Votes: 0GitHub stars: 3
- Avdner EvalEvaluates the ability of audio-visual models to separate cinematic audio into speech, music, and sound effects using visual cues like lip movements and scene context. It probes cross-track isolation, perceptual fidelity, and the model's capacity to leverage multi-stream video information for source disentanglement. Use when the user wants to benchmark on AVDnR, or asks about evaluating this task. Reports FAD.Votes: 0GitHub stars: 3
- Avagent EvalEvaluates the quality of audio-visual joint representations by testing downstream capabilities including classification, sound source localization, segmentation, and source separation. It measures how well an agentic workflow aligns audio and video modalities to improve cross-modal recognition and spatial/temporal synchronization. Use when the user wants to benchmark on VGGSound-Music, VGGSound-Instruments, MUSIC, Flickr-SoundNet, AVSBench, VGGSound-All, AudioSet, or asks about evaluating thi...Votes: 0GitHub stars: 3
- Av Speech Separation EvalEvaluates a model's ability to separate target speaker speech from audio mixtures (noise or other speakers) using synchronized visual face cues. It probes speaker-independent audio-visual fusion and robustness to varying numbers of speakers and background noise. Use when the user wants to benchmark on AVSpeech, AudioSet, CHiME-2, Mandarin, TCD-TIMIT, CUAVE, or asks about evaluating this task. Reports SDR improvement.Votes: 0GitHub stars: 3
- Av Speech Enhancement EvalThis benchmark evaluates a model's ability to isolate a target speaker's voice from multi-talker audio environments using only lip-region video inputs. It probes audio-visual speech enhancement, testing how well the network predicts magnitude and phase masks to suppress interference and noise while preserving speech intelligibility and perceptual quality. Use when the user wants to benchmark on LRS2, VoxCeleb2, or asks about evaluating this task. Reports PESQ.Votes: 0GitHub stars: 3
- Av Speakerbench EvalThis benchmark probes fine-grained audiovisual reasoning in multimodal large language models, specifically requiring them to jointly determine who is speaking, what is being said, and when events occur within real-world video clips. It evaluates cross-modal fusion, temporal grounding, and speaker-centric perception through multiple-choice questions validated by human experts. Use when the user wants to benchmark on AV-SpeakerBench, or asks about evaluating this task. Reports MCQ accuracy.Votes: 0GitHub stars: 3
- Av Odyssey Bench EvalEvaluates multimodal large language models' ability to perceive, integrate, and reason over interleaved audio and visual inputs. It probes basic auditory perception (e.g., loudness, pitch, duration) and complex cross-modal tasks spanning timbre, tone, melody, spatial reasoning, temporal dynamics, hallucination detection, and intricate reasoning across 10 domains. Use when the user wants to benchmark on AV-Odyssey Bench, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Av Deepfake1m EvalEvaluates models on detecting and temporally localizing audio-visual deepfakes in realistic, LLM-generated content. It probes robustness against multimodal manipulations like face reenactment and text-to-speech, testing both video-level classification and frame/segment-level localization. Use when the user wants to benchmark on AV-Deepfake1M, or asks about evaluating this task. Reports AP@0.5.Votes: 0GitHub stars: 3