Skip to content
Back to skills

Numerical Optimization

ASecurity

Practical optimization — choosing algorithms, handling constraints, tuning, and diagnosing convergence failures.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 29, 2026
ai-agentspythonrustgodebugging

Works with

  • cli

Security analysis

A100/100

Scanned September 29, 2026

npx -y skills add aicodedecode/awesome-muse-skills --skill numerical-optimization --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Numerical Optimization?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Numerical Optimization
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-numerical-optimization/badge)](https://www.skillsdirectory.com/skills/aicodedecode-numerical-optimization)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: numerical-optimization
description: Practical optimization — choosing algorithms, handling constraints, tuning, and diagnosing convergence failures.
category: scientific
---

## Overview

Optimization is everywhere in science: fitting models, designing
experiments, training ML, calibrating simulations. This skill covers
selecting the right algorithm for the problem structure (smooth vs
non-smooth, constrained vs free, convex vs non-convex), setting up
objectives correctly (scaling, gradients), and diagnosing the failures
that waste weeks of compute.

## When to use

- Fitting a model to data (least squares, maximum likelihood)
- Choosing between gradient-based, derivative-free, and global optimizers
- Handling constraints: bounds, equalities, inequalities
- Tuning hyperparameters or calibrating simulation parameters
- Debugging an optimizer that stalls, diverges, or returns nonsense

## Core concepts

- **Problem taxonomy:** convex (one global minimum — use gradient methods confidently) vs non-convex (many minima — need global search or multi-start); smooth (gradients help) vs non-smooth/noisy (derivative-free: Nelder–Mead, CMA-ES, Bayesian optimization).
- **Gradient methods:** steepest descent (slow), conjugate gradient, (L-)BFGS (quasi-Newton, the default for smooth unconstrained problems), Newton (needs Hessians). Supply analytic or autodiff gradients — finite differences are slower and noisier.
- **Constrained optimization:** bounds (L-BFGS-B), general constraints (SLSQP, trust-constr, augmented Lagrangian, interior-point). Reformulate when possible: log-transforms turn positivity into unconstrained problems.
- **Global optimization:** basin-hopping, differential evolution, dual annealing, Bayesian optimization (expensive objectives) — for rugged landscapes; always follow with a local polish.
- **Scaling:** optimizers assume ~O(1) variables; rescale parameters to similar magnitudes or the Hessian conditioning destroys convergence.
- **Least squares structure:** separable linear/nonlinear parameters (variable projection), sparse Jacobians — exploit structure instead of treating everything as a black box.

- **Automatic differentiation:** exact gradients at the cost of a function evaluation (JAX, autograd) — eliminates finite-difference noise entirely; the single biggest upgrade for gradient-based optimization of coded objectives.
- **Stochastic optimization:** SGD and Adam for noisy/large-scale objectives (machine learning) — the noise is a feature (escapes sharp minima); learning-rate schedules and batch sizes are the tuning knobs.
- **Derivative-free for the truly black box:** Bayesian optimization builds a surrogate (Gaussian process) and explores efficiently — the method of choice when each evaluation costs minutes or dollars.

## Practical workflow

### 1. Formulate well

```python
from scipy.optimize import minimize
# Scale variables to O(1); supply a gradient; bound what must stay physical
res = minimize(fun, x0_scaled, jac=grad, method="L-BFGS-B",
               bounds=bounds, options={"ftol": 1e-10, "maxiter": 10000})
```

1. Define the objective precisely: what is minimized, over what variables, subject to what — write it mathematically before coding.
2. Scale all variables to order unity; log-transform strictly positive parameters.
3. Provide gradients (analytic or autodiff via JAX/autograd) — finite-difference gradients fail on noisy objectives.

### 2. Choose the algorithm

| Problem | First choice |
|---|---|
| Smooth, unconstrained | L-BFGS (with gradients) |
| Smooth, bounds | L-BFGS-B |
| Smooth, general constraints | trust-constr / SLSQP |
| Non-smooth, low-dim | Nelder–Mead |
| Noisy/expensive, low-dim | Bayesian optimization |
| Non-convex, global search | Differential evolution → local polish |
| Least squares | Levenberg–Marquardt / trust-region reflective |

### 3. Multi-start and validate

1. For non-convex problems: multi-start from diverse initial points (Latin hypercube) — the best of many local optima beats one hopeful run.
2. Check convergence: gradient norm small, constraints satisfied, objective stable across restarts.
3. Validate the solution: does it make physical sense? Perturb the data slightly — stable solutions survive, overfit ones don't.

### 4. Diagnose failures

1. **Stalls immediately:** scaling (variables differ by 10⁶), bad gradient, or starting at a bound — check and rescale.
2. **Hits maxiter without converging:** loosen tolerances for noisy objectives, or the landscape is flat — reparameterize.
3. **Different answers per run:** non-convex landscape or stochastic objective — go global or multi-start; fix random seeds for reproducibility.
4. **Constraint violations:** penalty weights too small, or the feasible set is empty — check feasibility first.

### 5. Optimize an expensive black-box function

1. Define tight bounds from physical plausibility — Bayesian optimization wastes evaluations exploring absurd regions otherwise.
2. Choose the surrogate and acquisition function (expected improvement is the default); seed with a space-filling design (Latin hypercube, ~10 points per dimension).
3. Run in batches where parallel evaluation is possible; stop on a budget (evaluations or time), not on a tolerance — expensive objectives never get tight tolerances.

## Common pitfalls

- **Unscaled variables:** the #1 convergence killer — a parameter in nanometers next to one in gigapascals breaks every quasi-Newton method.
- **Finite-difference gradients on noisy objectives:** noise gets amplified into garbage search directions — use autodiff or derivative-free methods.
- **Single start on non-convex problems:** finding a local minimum and declaring victory — multi-start is mandatory.
- **Over-tight tolerances:** demanding 1e-12 on a noisy experimental objective wastes iterations chasing noise.
- **Ignoring constraints in the formulation:** optimizing then clipping to bounds gives infeasible "optima" — constrain properly.
- **Black-boxing structured problems:** least-squares and separable problems have specialized solvers 10–100× faster than generic ones.
- **Optimizing a stochastic objective as if deterministic:** re-evaluating the same point gives different values — use noise-aware methods (replicated evaluations, stochastic kriging) or the optimizer chases noise.
- **Ignoring integer/discrete variables:** rounding continuous optima to integers can be arbitrarily bad — use mixed-integer methods (or enumerate when the discrete space is small).

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…