Data & Analytics
Data analysis, BI, visualization, datasets, statistics, and ML workflows
Browse data & analytics skills
Showing 6,673–6,696 of 13,285 skills
Use this skill when you want to build a cleaned and deduplicated open text corpus from RedPajama for LLM pretraining. Avoid it when you already have a sufficient clean text corpus.
Use this skill when you want to remove semantically redundant examples from large-scale training data using embedding-based deduplication. Avoid it when you only need exact or near-exact duplicate removal or your dataset is small enough that redundancy is not a concern.
Use this skill when you want to evaluate and reduce object hallucination in VLMs using a polling-based evaluation framework. Avoid it when you are not concerned about object hallucination or have other evaluation methods.
Use this skill when you want to use ImageNet-21K for vision encoder pretraining with proper preprocessing and training recipes. Avoid it when ImageNet-1K or CLIP pretraining is sufficient.
Use this skill when you want to apply exact and near-duplicate removal to LLM training data using suffix arrays and MinHash. Avoid it when your dataset is already deduplicated or too small for dedup to matter.
Use this skill when you need a multilingual image-text dataset mined from Wikipedia covering 108 languages with curated metadata. Avoid it when you only need English data or web-crawled alt-text.
Use this skill when you want to construct a large diverse text corpus from 22 curated sources for LLM pretraining that is also used as the text component in multimodal training. Avoid it when you only need image-text data without a standalone text corpus.
Use this skill when you want to curate image-caption data from Reddit with transparent provenance and community-driven quality. Avoid it when you need professionally written captions or cannot comply with Reddit's terms of use.
Use this skill when you need a large-scale detection dataset with 365 categories for training open-vocabulary detection models used in VLM pipelines. Avoid it when you only need the 80 COCO categories or a smaller detection dataset.
Use this skill when you need a benchmark testing mathematical reasoning in visual contexts including charts, plots, diagrams, and geometry. Avoid it when you do not need mathematical reasoning evaluation.
Use this skill when you need a dataset linking noun phrases in captions to bounding box regions in images for visual grounding. Avoid it when you do not need phrase-level grounding or the 31K image scale is too small.
Use this skill when you need a massive egocentric video dataset with diverse annotations for training video-language models. Avoid it when you do not work with egocentric/first-person video or need third-person video data.
Use this skill when you want to build an image captioning dataset from web alt-text using automated cleaning and hypernym replacement. Avoid it when you need fine-grained or domain-specific captions that web alt-text cannot provide.
Use this skill when you want to use Whisper ASR to transcribe speech from videos for creating video-text training data. Avoid it when your videos do not have speech or you have existing transcripts.
bu skill covers Dağıt:ing anomaly Tespit systems for industrial control environments using machine learning models trained on OT network baselines, physics-based process models, and behavioral
W&B expert: experiment tracking, hyperparameter search, artifact management, sweep, team dashboards, performance visualization. Use when tracking ML experiments with Weights & Biases. Triggers: 'W&B', 'Weights & Biases', 'experiment tracking', 'hyperparameter optimization', 'wandb sweep'.
Expert-level Statistician skill covering frequentist and Bayesian statistical analysis, experimental design, causal inference, survival analysis, mixed models, multiple testing correction, and statistical consulting. Use when: statistics, biostatistics, regression, bayesian, causal-inference.
Expert SPSS and SAS user for statistical analysis. Use when running descriptive statistics, hypothesis tests, regression models, or survey analysis
R statistics expert: tidyverse, ggplot2, statistical modeling, hypothesis testing, regression analysis. Use when doing statistical analysis, data visualization, or predictive modeling with R.
Elite public health analyst specializing in epidemiological surveillance, health policy analysis, program evaluation, and population health assessment. Transforms health data into evidence-based recommendations for community health.
Pandas expert: DataFrame operations, merge/join, groupby, time series, performance optimization. Use when analyzing data, building ETL pipelines, or data manipulation with Python.
Expert Looker and Metabase user for business intelligence and embedded analytics. Use when building dashboards, creating data models, or implementing self-service analytics
Expert skill for linkedin-engineer
Expert digital twin architect with 10+ years designing cyber-physical systems for manufacturing, infrastructure, and smart cities. Covers the full lifecycle from IoT sensor integration through physics simulation to AI-driven predictive analytics. Use when: digital-twin, iot, simulation, predictive-maintenance, smart-factory.