Data & Analytics
Data analysis, BI, visualization, datasets, statistics, and ML workflows
Browse data & analytics skills
Showing 3,913–3,936 of 13,076 skills
Wasserstein Exponential Smoothing methodology from arXiv:2606.05560 — extends classical exponential smoothing to distributional time series in Wasserstein space. Provides consistent parameter estimation via Wasserstein distance minimization, applicable to high-frequency financial returns, electricity demand, and any distribution-valued time series forecasting. Activation: wasserstein exponential smoothing, distributional time series, Wasserstein forecasting, distributional forecasting, 分布时间序列...
Systematic RSA comparison showing untrained CNNs match backpropagation-trained networks at V1 cortical representations. Demonstrates that architectural biases alone (without learning) produce brain-aligned representations in early visual cortex, questioning the necessity of task-driven training for V1 modeling. (arXiv:2604.16875, April 2026)
Systematic RSA comparison showing untrained CNNs match backpropagation-trained CNNs at V1 visual cortex. Use for understanding representational similarity, visual cortex modeling, and the role of training in neural network-brain alignment. Keywords: CNN, V1, RSA, untrained networks, backpropagation, representational similarity analysis, visual cortex.
Systematic RSA comparison showing untrained CNNs match backpropagation at V1 alignment with human fMRI. Evaluates BP, FA, PC, and STDP learning rules against THINGS-fMRI dataset using 720 stimuli across 3 subjects. Use when studying brain-model alignment, comparing learning rules, or analyzing visual cortex representations via Representational Similarity Analysis.
Systematic RSA comparison showing untrained CNNs match backpropagation-trained CNNs at V1 visual cortex. Trigger words: untrained CNN, backpropagation, RSA, V1, representational similarity
Systematic RSA comparison showing untrained CNNs match backpropagation-trained networks at V1 visual cortex, revealing architecture's dominant role over learning rules in neural alignment. Activation triggers: untrained cnn, backpropagation, v1, rsa, representational similarity, learning rules, architecture-driven.
Systematic RSA comparison showing untrained CNNs match backpropagation-trained CNNs in V1 visual cortex alignment. Large-scale fMRI analysis reveals that random feature detectors can capture V1 representational structure. Keywords: untrained CNN, V1 cortex, backpropagation, RSA, representational similarity, visual cortex, fMRI.
**arXiv ID:** 2605.20009 **Authors:** Vy Bui, Hang Yu, Karthik Kantipudi, Ziv Yaniv, Stefan Jaeger **Published:** 2026-05-19T15:39:36Z **Abstract:** Backpropagation with gradient descent is a common optimization strategy employed by most neural network architectures in machine learning. However, finding optimal hyperparameters to guide training has proven challenging. While it is widely acknowledged that selecting appropriate parameters is crucial for avoiding overfitting and achieving unbias...
**arXiv ID:** 1806.11379 **Authors:** Tomaso Poggio, Qianli Liao, Brando Miranda, Andrzej Banburski, Xavier Boix, Jack Hidary **Published:** 2018-06-29T12:39:08Z **Abstract:** A main puzzle of deep neural networks (DNNs) revolves around the apparent absence of "overfitting", defined in this paper as follows: the expected error does not get worse when increasing the number of neurons or of iterations of gradient descent. This is surprising because of the large capacity demonstrated by DNNs to ...
First systematic application of Successor Representations (SRs) from reinforcement learning to natural language. Trains deep residual network on WikiText-103 to predict future word distributions; structured language representations (noun/verb/adjective categories) emerge spontaneously without explicit linguistic supervision. Establishes bridge between RL, linguistics, and cognitive neuroscience. Based on arXiv:2605.24585 (May 2026). Use when studying successor representations in language, eme...
Analysis of Neural Tangent Kernel (NTK) collapse near dynamical bifurcations in state-space models. Studies how the NTK spectrum degrades as recurrent networks approach critical transitions. Activation: NTK collapse, bifurcation analysis, state-space NTK, critical transitions neural networks, dynamical systems deep learning.
信噪比和样本数量调控神经网络表征对齐的方法论。研究神经网络潜在表征的通用性规律,揭示对齐与数据质量和数量的非平凡依赖关系。适用于表征对齐分析、神经网络可解释性、训练优化。触发词:表征对齐、SNR、样本数量、插值阈值、通用表征。
**arXiv ID:** 2410.16151 **Authors:** Mostafa Hussien, Mahmoud Afifi, Kim Khoa Nguyen, Mohamed Cheriet **Published:** 2024-10-21T16:18:31Z **Abstract:** Recent advancements have scaled neural networks to unprecedented sizes, achieving remarkable performance across a wide range of tasks. However, deploying these large-scale models on resource-constrained devices poses significant challenges due to substantial storage and computational requirements. Neural network pruning has emerged as an effe...
Formalizes how attractor networks emerge from the free energy principle applied to universal partitioning of random dynamical systems. Results in self-orthogonalizing attractor representations, biologically plausible multi-level Bayesian active inference. Use when: studying attractor dynamics in neural networks, free energy principle applications, Bayesian active inference models, biologically plausible learning, Boltzmann Machine variants, self-organizing neural dynamics. Triggered by: free ...
SAE 最优性结构理论 - 解释 Sparse Autoencoders 如何从最优性条件提取可解释特征。涵盖层次分裂与吸收、残差结构、密集对立特征等现象的理论基础。
**arXiv ID:** 2404.03992 **Authors:** Mohammed Ghaith Altarabichi, Sławomir Nowaczyk, Sepideh Pashami, Peyman Sheikholharam Mashhadi, Julia Handl **Published:** 2024-04-05T10:02:32Z **Abstract:** This paper investigates how various randomization techniques impact Deep Neural Networks (DNNs). Randomization, like weight noise and dropout, aids in reducing overfitting and enhancing generalization, but their interactions are poorly understood. The study categorizes randomness techniques into four...
**arXiv ID:** 2009.06202 **Authors:** Johannes Lederer **Published:** 2020-09-14T05:06:59Z **Abstract:** It has been observed that certain loss functions can render deep-learning pipelines robust against flaws in the data. In this paper, we support these empirical findings with statistical theory. We especially show that empirical-risk minimization with unbounded, Lipschitz-continuous loss functions, such as the least-absolute deviation loss, Huber loss, Cauchy loss, and Tukey's biweight loss...
Renormalization group (RG) framework for analyzing scaling laws and criticality in brain activity. Connects 1/f noise, neuronal avalanches, and coarse-grained descriptions through RG theory. Activates: renormalization brain, scaling law neural activity, 1/f noise brain, neuronal avalanche scaling, coarse-graining neural dynamics, RG criticality brain, power law neural scaling.
**arXiv ID:** 1806.02997 **Authors:** Aleksei Vasilev, Vladimir Golkov, Marc Meissner, Ilona Lipp, Eleonora Sgarlata, Valentina Tomassini, Derek K. Jones, Daniel Cremers **Published:** 2018-06-08T07:28:36Z **Abstract:** In machine learning, novelty detection is the task of identifying novel unseen data. During training, only samples from the normal class are available. Test samples are classified as normal or abnormal by assignment of a novelty score. Here we propose novelty detection methods...
**arXiv ID:** 2412.17411 **Authors:** Jeonghwan Cheon, Se-Bum Paik **Published:** 2024-12-23T09:22:00Z **Abstract:** Uncertainty calibration is crucial for various machine learning applications, yet it remains challenging. Many models exhibit hallucinations - confident yet inaccurate responses - due to miscalibrated confidence. Here, we show that the common practice of random initialization in deep learning, often considered a standard technique, is an underlying cause of this miscalibration,...
Preisach Attention Layer (PAL) — a novel sequence modeling architecture that replaces softmax attention with the classical Preisach hysteresis operator from mathematical physics. Uses binary relay operators with learned thresholds and a stack of local extrema as internal state. Achieves Turing-completeness at O(1) depth via two-stack PDA simulation. Activation: attention, hysteresis, sequence modeling, episodic memory, transformer alternative, rate-independent computation
**arXiv ID:** 2506.00920 **Authors:** Philip Heejun Lee **Published:** 2025-06-01T09:20:44Z **Abstract:** Deep sequence models typically degrade in accuracy when test sequences significantly exceed their training lengths, yet many critical tasks--such as algorithmic reasoning, multi-step arithmetic, and compositional generalization--require robust length extrapolation. We introduce PRISM, a Probabilistic Relative-position Implicit Superposition Model, a novel positional encoding mechanism tha...
Physics-guided neural network design and training methods. Embed physical laws, constraints, and symmetries into neural network architecture for improved modeling of physical systems (quantum mechanics, statistical physics, fluid dynamics, materials science). Activation: physics guided neural network, 物理学指导神经网络, physics-informed neural network, PINN, physics-constrained learning, quantum neural network, physics-aware training.
On-Policy Distillation (OPD) methodology for transforming autoregressive models into diffusion language models efficiently, eliminating train-inference mismatch.