Category

Research

Research, evidence gathering, literature, reports, investigation, and synthesis

20,834
skills in category
869
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse research skills

Showing 6,505–6,528 of 20,834 skills

Lexical Simplification EvalA

Evaluates the robustness and classification accuracy of pre-trained language models when augmented with rule-based lexical simplification as auxiliary inputs. It probes whether lemmatization and rare-word replacement preserve semantic meaning while mitigating lexical diversity effects on downstream NLU tasks. Use when the user wants to benchmark on SST-2, CR, SUBJ, MR, AG, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Lexam EvalA

Probes large language models' ability to perform structured, multi-step legal reasoning on real-world law exam questions. It evaluates both open-ended legal analysis and multiple-choice selection across diverse jurisdictions and legal domains. Use when the user wants to benchmark on LEXam, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Lex Bench EvalA

Evaluates text-to-image generation models on their ability to accurately render specified text within images, control visual attributes (color, position, font), and maintain aesthetic quality. It measures OCR fidelity, attribute controllability, and human-perceived aesthetics. Use when the user wants to benchmark on LeX-Bench, SimpleBench, CreateBench, AnyText-Benchmark, or asks about evaluating this task. Reports PNED.

researchpythongo
0
3
Lewmm Physical EvalA

Evaluates whether a latent world model captures physical structure and dynamics by probing latent representations for physical quantities and measuring predictive surprise under physical versus visual perturbations. Use when the user wants to benchmark on TwoRoom, PushT, OGBench-Cube, Reacher, or asks about evaluating this task. Reports MSE.

researchpythongo
0
3
Leum Vl EvalA

Evaluates a video-language model's capacity for structured video understanding across six professional dimensions (subject, aesthetics, camera language, editing, narrative, dissemination) while maintaining general multimodal capabilities. It probes timeline-grounded reasoning, temporal localization, and document/OCR comprehension. Use when the user wants to benchmark on FeedBench, Open Benchmarks (Video-MME, MVBench, MMBench-EN, etc.), or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Lemur Retrieval EvalA

Evaluates the ability of multilingual embedding models to retrieve relevant legislative documents given structured metadata queries. It probes cross-lingual semantic alignment and domain-adaptive retrieval performance across varying language resource levels. Use when the user wants to benchmark on LEMUR, or asks about evaluating this task. Reports Acc@k.

researchpythongo
0
3
Lempel Ziv Complexity And SensitivityA

Evaluates how neural network hyperparameters (activation functions, depth, learning rate) affect output complexity and robustness to input perturbations. Use when the user has predictions and gold and needs to compute Lempel-Ziv Complexity.

researchpythongo
0
3
Lemat Bulk EvalA

Evaluates the accuracy and robustness of crystal structure fingerprinting and hashing algorithms for de-duplicating quantum chemistry materials databases. It probes sensitivity to structural perturbations (atomic noise, lattice strain, translations) and performance on disordered crystal systems. Use when the user wants to benchmark on LeMat-Bulk, or asks about evaluating this task. Reports success rate.

researchpythongo
0
3
Lemas Multilingual Tts Edit EvalA

This benchmark evaluates multilingual text-to-speech synthesis and text-based speech editing capabilities. It probes pronunciation stability, cross-lingual generalization, and the perceptual naturalness of localized audio edits across multiple languages. Use when the user wants to benchmark on LEMAS-Dataset, or asks about evaluating this task. Reports WER.

researchpython
0
3
Lego Egocentric Action EvalA

Evaluates a diffusion model's ability to generate egocentric action frames from a pre-action image and a text prompt. It probes the model's capacity to capture action state transitions while preserving contextual information and aligning with natural language instructions in egocentric video domains. Use when the user wants to benchmark on Ego4D, Epic-Kitchens-100, or asks about evaluating this task. Reports EgoVLP score, EgoVLP+ score.

researchpythongo
0
3
Legalbenchpt EvalA

Evaluates large language models' ability to reason about and classify Portuguese legal concepts across 31 distinct legal domains. It probes zero-shot question-answering capabilities using multiple-choice, true/false, matching, and case-analysis formats derived from law exam questions. Use when the user wants to benchmark on LegalBench.PT, or asks about evaluating this task. Reports balanced accuracy.

researchpythongo
0
3
Legalbench Rag EvalA

Evaluates the retrieval fidelity of RAG systems in the legal domain by measuring how precisely and completely a model retrieves minimal, highly relevant text snippets from legal documents to answer specific queries. Use when the user wants to benchmark on LegalBench-RAG, or asks about evaluating this task. Reports Precision.

researchpythongo
0
3
Legalbench EvalA

This benchmark probes large language models' ability to perform diverse, real-world legal reasoning tasks, including rule-recall, issue-spotting, rule-application, interpretation, and rhetorical understanding. It evaluates how well models can apply legal frameworks, classify contractual clauses, and answer questions based on statutory or case law text. Use when the user wants to benchmark on LegalBench, or asks about evaluating this task. Reports accuracy.

researchpythongo
0
3
Legal Zero Days EvalA

This benchmark probes an AI system's ability to detect previously undiscovered legal vulnerabilities within governance frameworks. It tests whether models can identify systemic flaws that could cause immediate disruption without requiring traditional litigation, measuring their capacity for advanced legal reasoning and regulatory logic parsing. Use when the user wants to benchmark on Legal Zero-Days, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Legal Text Classification EvalA

This benchmark assesses models on hierarchical legal text classification tasks (LAP, JP, CP) across three Swiss languages. It tests domain-specific classification accuracy and robustness to long legal documents and multilingual inputs. Use when the user wants to benchmark on Legal Classification (LAP, JP, CP), or asks about evaluating this task. Reports Hierarchical Macro-F1.

researchpythongo
0
3
Legal Ner EvalA

Evaluates a model's capacity to identify and classify 14 domain-specific legal entities (e.g., Court, Statute, Precedent, Petitioner Name) within unstructured legal documents. This probes fine-grained information extraction capabilities tailored to legal terminology and structure. Use when the user wants to benchmark on LegalEval L-NER Dataset, or asks about evaluating this task. Reports standard F1 score.

researchpythongo
0
3
Legal Lens EvalA

This benchmark evaluates language models on identifying legal violations and associating victims in unstructured legal text. It probes two core capabilities: named entity recognition for specific causes of action and natural language inference for linking victims to legal claims across different legal domains. Use when the user wants to benchmark on LegalLens NER Dataset, LegalLens NLI Dataset, or asks about evaluating this task. Reports Macro F1.

researchpythongo
0
3
Legal Information Retrieval EvalA

This benchmark evaluates information retrieval systems on Swiss legal rulings and legislation. It tests the ability to rank relevant legal documents against long, multilingual queries and corpora. Use when the user wants to benchmark on Legal Information Retrieval, or asks about evaluating this task. Reports NDCG.

researchpythongo
0
3
Legal Constitution EvalA

Evaluates a fine-tuned open-source language model on its ability to perform keyword extraction, summarization, and sentiment analysis on the Indian Constitution. The protocol tests whether domain-specific fine-tuning improves the model's capacity to grasp nuanced legal semantics and structural elements. Use when the user wants to benchmark on Indian Constitution, or asks about evaluating this task. Reports precision, recall, and F1 score.

researchpythongo
0
3
Legal Cloze Test EvalA

This benchmark probes a model's ability to understand and predict precise legal terminology and procedural concepts in Turkish court documents. It evaluates both masked language modeling capabilities on legal cloze sentences and structural segmentation accuracy for parsing document sections. Use when the user wants to benchmark on Legal Cloze Test benchmark, v12 Court Decision Segmentation Dataset, or asks about evaluating this task. Reports Top-1 Accuracy.

researchpythongo
0
3
Legal Bert EvalA

Evaluates domain-adapted BERT models on legal text classification and named entity recognition to measure the impact of further pre-training and hyperparameter tuning strategies. Use when the user wants to benchmark on EURLEX57K, ECHR-CASES, CONTRACTS-NER, or asks about evaluating this task. Reports accuracy, F1.

researchpythongo
0
3
Legal Bench EvalA

Evaluates the ability of state-space models (Mamba/SSD-Mamba) and transformers to perform statutory classification and case law retrieval on long-context legal documents. It probes how well models capture fine-grained semantic distinctions and maintain global coherence over thousands of tokens while balancing accuracy with computational throughput. Use when the user wants to benchmark on SCOTUS, ILDC, ECtHR, EUR-Lex, or asks about evaluating this task. Reports Accuracy.

researchpythongo
0
3
Leetcode Dataset EvalA

Evaluates code generation and algorithmic reasoning capabilities on competitive programming problems. It specifically probes temporal robustness by testing on problems released after a strict cutoff date to detect data contamination, and measures performance across difficulty levels and algorithmic topics. Use when the user wants to benchmark on LeetCodeDataset, or asks about evaluating this task. Reports pass@1.

researchpythongo
0
3
Lecavrdv2 EvalA

Probes a model's ability to retrieve relevant Chinese criminal case documents from a large corpus based on legal queries. It specifically tests alignment with multi-dimensional legal relevance criteria, including case characterization, penalty matching, and procedural similarity. Use when the user wants to benchmark on LeCaRDv2, or asks about evaluating this task. Reports Recall@K.

researchpythongo
0
3