Category

Data & Analytics

Data analysis, BI, visualization, datasets, statistics, and ML workflows

13,078
skills in category
545
pages available
Security grades appear on each card once the skill has been scanned. Newly imported skills may briefly show without a grade until the backfill job runs.
Open in full browser

Browse data & analytics skills

Showing 4,921–4,944 of 13,078 skills

Justrl Simple RecipeA

Demonstrate that simple single-stage RL with fixed hyperparameters matches complex multi-stage approaches for training small LLMs on mathematical reasoning. Use basic setup: GRPO algorithm, rule-based verification, 16K token context, standard training data without difficulty filtering. Train two 1.5B models to competitive performance using 2× less compute than sophisticated approaches.

datagoperformance
0
6
Higher Order Linear Attention MechanismA

Enable data-dependent higher-order interactions in attention using prefix-sufficient statistics that maintain linear time and constant state, replacing quadratic dot-product attention while preserving expressivity through compact matrix operations.

datapythonexpress
0
6
High Entropy Minority Tokens RlA

Optimize only high-entropy tokens during RL training to achieve better reasoning performance with 80% fewer gradient updates.

datapythongo
0
6
Gradient Grouping Learning Rate ScalingA

Improve adaptive learning rates by clustering gradient statistics within layers and applying cluster-specific scaling.

datapython
0
6
Expert Threshold RoutingA

Improve MoE language model efficiency with causal threshold-based routing that eliminates auxiliary losses and enables dynamic per-token computation.

datapythongit
0
6
Exgrpo Experience Replay ReasoningA

Improve RLVR training efficiency by selectively replaying trajectories based on correctness and entropy. Medium-difficulty questions and low-entropy solutions are most valuable; selective replay yields +3.5-7.6% improvements.

datapythongit
0
6
Dcpo Dynamic Clipping Policy OptimizationA

DCPO eliminates zero-gradient dead zones in policy optimization by adaptively adjusting token-level clipping bounds based on prior probabilities and smoothing advantage standardization across cumulative training steps, achieving 28% improvement in effective response utilization and 10x reduction in token clipping ratio on mathematical reasoning benchmarks.

datapythongit
0
6
Config Knowledge DistillationA

Improve student model robustness under covariate shift by using diffusion-based augmentation that targets spurious features via teacher-student disagreement.

datapythongo
0
6
Bro Rl Broad Rollout ScalingA

Overcome reasoning model training plateaus by increasing rollouts per prompt (N=512) rather than training steps, addressing unsampled coupling that destabilizes learning. Theoretical analysis shows broad exploration eliminates plateau bottleneck.

datapythongo
0
6
Blockwise Advantage EstimationA

Improve credit assignment in multi-objective RL by decomposing advantages into segment-specific values. Use Outcome-Conditioned Baselines to reduce cross-objective interference without expensive rollouts, enabling better training signals for multi-step completions with different reward functions per segment.

datapython
0
6
Bifrost Patch Clip DiffusionA

Connects multimodal language models with diffusion models using patch-level CLIP embeddings as shared latent variables, enabling controllable image generation with minimal training overhead.

datapythonperformance
0
6
Ares Entropy ShapingA

Calibrate exploration effort in reasoning traces based on problem difficulty by detecting high-entropy windows and applying hierarchical entropy rewards. Reduces unnecessary reasoning on easy tasks while increasing exploration on hard tasks.

datapythongit
0
6
Alchemist Meta Gradient SelectionA

Select optimal training subsets for T2I models through meta-gradient-based rater networks. Score each sample based on gradient influence on validation performance without retraining. Implement shift-Gaussian pruning excluding high-scoring samples. Achieve 5× training speedup with 50% subset outperforming full dataset.

dataperformance
0
6
Agentic Llm Data ScienceA

Train agentic LLMs through curriculum-based learning to autonomously execute full data science workflows from raw data to analysis reports, enabling 8B models to match proprietary systems.

datapythonsql
0
6
Adaptive Data Refinement Vlm Long TailA

Mitigate long-tail distribution problems in VLM training data through adaptive rebalancing and diffusion-based synthesis. Uses entity distribution analysis to identify head/tail imbalance and applies targeted data augmentation, improving LLaVA 1.5 performance by 4.36% without increasing training data volume.

datapythongo
0
6
Wave TheoryA

Ocean wave theory including wave spectra, statistics, irregular seas, and wave transformation for offshore engineering

datapythongo
0
17
Ship Dynamics 6dofA

6DOF ship dynamics, equations of motion, seakeeping analysis, and natural frequency calculations

datapythongo
0
17
Risk AssessmentA

```yaml name: risk-assessment version: 1.0.0 category: sme tags: [risk, probabilistic, monte-carlo, reliability, uncertainty, sensitivity-analysis, marine-safety] created: 2026-01-06 updated: 2026-01-06 author: Claude description: | Expert risk assessment and probabilistic analysis for marine and offshore operations. Includes Monte Carlo simulations, reliability calculations, uncertainty quantification, and decision-making under risk. ```

datapythongo
0
17
Orcaflex SpecialistA

```yaml name: orcaflex-specialist version: 1.0.0 category: sme tags: [orcaflex, offshore, simulation, python-api, automation, marine-dynamics, mooring, riser] created: 2026-01-06 updated: 2026-01-06 author: Claude description: | Expert OrcaFlex workflows, Python API automation, model validation, and best practices for offshore marine simulations. Covers mooring analysis, riser dynamics, installation simulations, and advanced post-processing. ```

datapythongo
0
17
Pandas Data ProcessingA

Pandas for time series analysis, OrcaFlex results processing, and marine engineering data workflows

datapythongo
0
17
Numpy Numerical AnalysisA

NumPy for matrix operations, FFT, linear algebra, and numerical computations in marine engineering

datapythongo
0
17
Ydata ProfilingA

Automated data quality reports with comprehensive variable analysis, missing value detection, correlations, and HTML report generation - formerly pandas-profiling

datapythongo
0
17
SweetvizA

Automated EDA comparison reports with target analysis, feature comparison, and HTML report generation for pandas DataFrames

datapythongo
0
17
StreamlitA

Build interactive data applications and dashboards with pure Python - no frontend experience required

datajavascriptpython
0
17