Data & Analytics
Data analysis, BI, visualization, datasets, statistics, and ML workflows
Browse data & analytics skills
Showing 10,345–10,368 of 13,069 skills
Extract TCR:pMHC specificity data from raw source files (papers, supplementary tables, XLS, PDF, 10X output, AIRR-format) and produce a VDJdb-formatted TSV chunk ready for /format and /proofread.
Identify and classify duplicate TCR records across all VDJdb chunks at three resolution levels (beta-only, paired, and same-epitope multi-MHC), categorise by publication source, author overlap, and flag spurious high-frequency records.
Forecasting models and the `ts_forecast_by` / `ts_forecast_var_by` API surface of the anofox_forecast DuckDB extension. Covers 36 models (baseline, exponential smoothing, state-space ARIMA + Kalman, classical GARCH, Theta, multi-seasonal, intermittent-demand, distributional Laplace with three variants, panel/global GlobalETS/GlobalTheta/GlobalCroston, and multivariate VAR via ts_forecast_var_by), plus exogenous-regressor forecasting (ARIMAX / ThetaX / MFLESX via ts_forecast_exog_by / ts_forec...
Exploratory data analysis, data quality, and statistical diagnostics for the anofox_forecast DuckDB extension — 34 per-series statistics, data-quality scoring, quality-report summaries, 117 tsfresh-compatible feature extraction, and 7 diagnostic functions covering stationarity (ADF, KPSS, combined verdict) and residual adequacy (Ljung-Box, Durbin-Watson, Jarque-Bera, combined report). Use before forecasting to understand series characteristics (length, gaps, trend, seasonality strength, inter...
Seasonality, changepoint, peak, and decomposition detection for the anofox_forecast DuckDB extension. Use when identifying seasonal periods before configuring seasonal forecasting models, detecting structural breaks, analysing peak timing regularity, or decomposing a series into trend / seasonal / residual components.
Data preparation for the anofox_forecast DuckDB extension — filling gaps, imputing nulls, dropping bad series, differencing, detrending, hierarchical key operations. Use when preparing raw time series for downstream forecasting or backtesting with `ts_forecast_by` / `ts_cv_folds_by`.
Backtesting, cross-validation, evaluation metrics, and conformal prediction intervals for the anofox_forecast DuckDB extension. Use when evaluating forecast accuracy, comparing models with time-series-aware CV, computing metrics (MAE / RMSE / MAPE / MASE / coverage), or attaching distribution-free prediction intervals to forecasts.
Use this to author and change a dbt project or a semantic layer: bootstrap a project in a repo that has none (`transform init`), write or refactor model SQL from staging to marts, add tests and docs in schema.yml, manage dependencies, and define or update the semantic layer, whether that is dbt semantic models (MetricFlow: entities, dimensions, measures, metrics) or native Apache Ossie documents in a repo with no dbt project at all. Reach for this rather than editing model files by hand whene...
Use this to keep a dbt project and its semantic layer correct as the warehouse and the business change, including a semantic layer that is native Apache Ossie documents rather than dbt. It detects drift on four axes and proposes the fix: schema drift (source columns and tables added, dropped, retyped, or renamed), volume drift (a row count that collapsed, a table that emptied, a load that half-failed), grain drift (a key that lost uniqueness, a changed row-per-entity cardinality, an increased...
Use this whenever you need to know what is actually in a database, warehouse, or DuckDB file before you trust it: ranked inventory of what exists, column profiles, PII detection, grain and data-quality problems, verified join inference, Mermaid ER diagrams, guarded ad-hoc SQL probes, k-means segmentation, and reading the semantic layer a repo declares (dbt semantic models, a hosted dbt Cloud layer, or native Apache Ossie documents), producing a draft map without dumping the whole schema into ...
Fill missing and refresh obsolete model descriptions in `packages/llm-info/data/models.yml` by querying OpenRouter and provider documentation. Use when the user asks to populate model descriptions, enrich the model catalog, or curate descriptions after running `pnpm sync-models`.
Work inside the user's live marimo notebook from the code editor: run Python in the same kernel the user does, inspect live notebook state, and commit durable notebook changes through code mode. Use whenever you create, analyze, or improve the user's marimo notebook.
Raw AdMapix ad creative data search. Use for 搜广告, 找素材, 广告视频, 创意素材, 竞品广告, ad creative, search ads, find creatives, competitor ads, ad spy. Returns structured JSON data only.
Turn a data file (CSV/TSV, Excel, SQLite, JSON/JSONL, Markdown, text, logs) into a static, shareable infographic IMAGE — a rendered PNG plus its editable HTML source. Composes real infographics (hero stat, numbered story spine, pictograms, annotations, takeaway), not styled dashboards. Use this skill whenever the user asks for an "infographic", "poster image", "one-pager", "social card", "an image I can share", "a PNG of this data", or any static picture that tells the data's story — as oppos...
Parse data files of many formats — CSV/TSV, Excel (.xlsx), SQLite, JSONL/NDJSON, JSON, Markdown, plain text, and log files — and generate a stunning, self-contained, single-file HTML experience that visualizes the data: a Three.js animated hero, GSAP scroll animations, Tailwind styling, column-profiled data tables with search/sort/pagination, histograms and time-series charts, a markdown reading view with outline navigation, a filterable log viewer, and a collapsible tree explorer for arbitra...
对中国上市公司相关的荐股文章、公众号推文、投资点评、雪球/小红书帖子、业绩说明会转述等"已存在的内容"做事实核验:把其中的财务数字与表态逐条拆出, 强制对齐 A 股/港股官方披露(巨潮资讯网 CNINFO、沪深北交易所、互动易/上证e互动、定期报告、临时公告、招股书),输出一份"哪些为真 / 哪些对不上 / 哪些查无此据 / 哪些是纯话术"的核验体检报告。当用户贴出一篇荐股文、股票点评、公司分析或截图并想知道"这靠不靠谱 / 数据是不是真的 / 帮我核实一下"时,当用户想核对某上市公司被引用的营收、净利润、毛利率、订单、市占率、股权或战略表态时,都要使用本 skill。 关键区分:本 skill 只"审计已有内容里的断言",方向是从内容出发、去官方披露里对账;它不"从 ticker 生成一份新的个股研究报告"—— 凡是"帮我分析这只股票 / 做估值 / 做杜邦 / 建个模型"这类生成与判断类需求,不属于本 skill 范围。
全流程问卷/量表数据分析方法学助手(中英文通用 / bilingual)。当用户需要分析问卷、量表或调查数据的任何环节时使用:量表与问卷设计、Likert 计分、数据清洗与缺失值/异常值处理、反向计分、描述性统计与正态性、共同方法偏差(Harman/CLF)、信度(Cronbach's α / CR 组合信度)、效度(内容/结构效度、EFA 探索性因子分析、CFA 验证性因子分析、收敛/区分效度 AVE)、差异检验(t检验、ANOVA、卡方、非参数)、相关与回归(含多重共线性 VIF)、结构方程模型 SEM、中介与调节效应(Bootstrap / PROCESS)、效应量、APA 表格与结果写作。工具中立,实操以 SPSS / AMOS 为主并在关键处给出 R(psych/lavaan/PROCESS) 做法。触发词包括"问卷分析""量表""信度效度""Cronbach""因子分析""EFA/CFA""SPSS""AMOS""结构方程""SEM""中介效应""调节效应""回归分析""方差分析""Likert""问卷数据"等。
老子道涨停股批量归因分析Skill。基于开源大模型,输入日期自动获取当日所有涨停股票数据,逐一输出多维度归因分析报告。GitHub https://github.com/laozdao/limit-up-dao。触发场景:用户要求分析某日涨停股、批量涨停归因、今日涨停复盘、涨停板原因分析、连板股分析等。
Surveys the current repo's open GitHub issues, ranks them by triage label and dependency graph, and recommends an optimal execution order; when you pick one to start, it gates on whether the issue is clear enough to execute and routes unclear ones to the grill-with-docs skill before any code is written. It can also render the board as a self-contained HTML map that groups issues into business lines, draws their dependency arrows, and spells out which feature each chain ships once completed. U...
의학연구 tabular 데이터(.xlsx/.csv)에 대해 한국어 단일 HTML EDA 리포트를 자동 생성하는 스킬. 행이 관찰 단위(환자·내원·병변·검체 등), 열이 변수인 모든 의학연구 데이터셋이 대상이며 연구 디자인(후향/전향 코호트, RCT·임상시험, case-control, cross-sectional, registry, survey 등)을 가리지 않는다. n·변수 타입별 요약, 결측 패턴, 분포 플롯, 이상치(implausible value) 감지, 선택적 소그룹별 Table 1, 상관관계 heatmap, VIF를 모두 한 파일에 임베딩한다. 사용자가 임상연구·관찰연구·임상시험·환자 데이터·registry·연구 데이터셋·엑셀/CSV 파일을 업로드하면서 "EDA", "데이터 탐색", "탐색적 분석", "기초통계", "Table 1", "결측 보고", "분포 확인", "데이터 살펴봐", "데이터 점검" 같은 표현을 사용하면 적극적으로 트리거하라. 단순 통계 분석(t-te...
의학연구 tabular 데이터(.xlsx/.csv)에서 baseline characteristics을 비교하는 Table 1을 자동 생성하는 스킬. RCT, prospective/retrospective cohort, case-control, cross-sectional, registry, single-arm 등 연구 디자인을 모두 지원하며 디자인에 따라 p-value 보고 정책(RCT는 CONSORT 2010에 따라 baseline p 숨김)이 자동 분기된다. 연속형은 정규성에 따라 mean±SD(Welch t/ANOVA)와 median[IQR](Mann-Whitney/Kruskal-Wallis)로, 범주형은 n(%)와 chi-square/Fisher's exact/Monte Carlo chi-square로 처리한다. 모든 변수에 대해 SMD(2군은 표준, ≥3군은 max pairwise)를 색 코딩(<0.1 ok / 0.1-0.2 small / ≥0.2 meaningful)...
Use when 用户请求选择、创建、改进或统一数据可视化,并需按数据与意图选择图表、主题及 D3.js、ECharts、Mapbox、Three.js、Matplotlib、Plotly 或 ggplot2 技术栈。
Lightweight, script-driven variant of data2motion for turning data into a smooth, on-brand animated chart (one self-contained HTML) — built for a TEXT-ONLY model that cannot see its own output. You never write HTML/CSS/SVG/animation; you extract the data, pick a chart, fill a small JSON spec, and run one build script that owns 100% of the look and the motion, so the result cannot drift from the house template style or lose its smoothness. Use to turn a number/stat/table/CSV/paragraph into a q...
一套模板驱动的数据可视化与报告生成 skill,既能严格从 Lupi、Basics、Glance、Maps 与 Interactive gallery 的真实实现生成 HTML 图表,也能从 12 套中英文整页报告模板生成可发布的 HTML 报告;以 Mono 为保底,能按数据语义自动选择内置彩色预设,也支持用户明确提供的自定义色板。地图仅在用户明确要求时启用,同一交付禁止混用色系。