Do AI Models Really Create?

I am a Ph.D. student exploring Embodied Intelligence. I write about Math, CS, AI and my views on them.
"Creativity is the ability to come up with ideas or artifacts that are new, surprising, and valuable" — Margaret A. Boden
Write your favorite number on a piece of paper. Is that number generated from you or just mixing of your experiences around you?
This seemingly trivial exercise opens a profound window into the nature of human creativity. When pressed to pick a "favorite" number, most people do not invoke pure randomness or abstract mathematical axioms in isolation. Instead, the choice emerges from a lifetime of culture conditioning, personal anecdotes, media exposure, and cognitive heuristics—perhaps a bias toward 7 (often cited as "lucky" in Western folklore), 42 (from The Hitchhiker's Guide to the Galaxy), or 37 (a frequent "random" pick in psychological studies). The distribution is far from uniform; it reflects a statistical recombination of learned patterns rather than spontaneous invention. Yet we intuitively feel the selection as *ours—*a personal style shaped by, but not reducible to, those inputs.
The same tension animates today's generative AI systems. Large language models, diffusion-based image generators, and reinforcement-learning agents produce outputs that feel strikingly original: poems never before written, images of surreal scenes never photographed, games moves that stun grandmasters. Are these genuine creations, or sophisticated recombinations drawn from vast training corpora? This blog examines the question through a dual lens—comparing human and AI learning pathways, formalizing both processes, designing and interpreting targeted experiments, and analyzing landmark AI "discoveries". The discussion draws on manifold learning, probabilistic generative modeling, and philosophical frameworks of creativity. By the end, a nuanced position emerges: current generative AI excels at high-dimensional recombination and exploratory navigation of learned manifolds, occasionally yielding functional novelty indistinguishable from creation, yet remains ontologically tethered to its data. True transformational invention—altering the conceptual space itself—remains rarer and hybrid.
Human Creativity
Humans learn by imitation from infancy. Infants mirror facial expressions; toddlers absorb language through statistical pattern extraction from parental speech; artists study masters before developing signature techniques. Cognitive science frames this as hierarchical Bayesian inference: the brain maintains internal generative models updated via prediction error (Friston's free-energy principle). Style arises when these models are conditioned on idiosyncratic experiences—emotional valence, cultural context, motor idiosyncracies—producing outputs that deviate systematically from the mean.
Consider number selection again. The histogram below simulates 5000 human choices with culturally amplified probabilities for cultural resonant values. The peaks and troughs illustrate bias, not uniform invention. Yet each individual's pick carries subjective ownership because the mixing occurs within a unique internal latent space shaped by personal history.
This is not merely copying; it is combinational creativity (Boden): familiar elements recombined in unfamiliar ways. Picasso absorbed African masks yet produced Cubism; the output transcended linear interpolation. Humans also perform exploratory creativity by navigating the rules of a domain (e.g., jazz improvisation within harmonic grammar) and, rarely, transformational creativity by breaking those rules (e.g., Schoenberg's twelve-tone system). Agency, emotion, and embodiment provide the "style" filter—something current AI approximates via conditioning or RLHF but lacks in qualia.
AI Learning and Generation
AI systems learn analogously but at superhuman scale and speed. Foundation models ingest trillions of tokens or images, optimizing parameters to approximate the data-generating distribution. Transformers capture long-range statistical dependencies; diffusion models learn to reverse noise addition; reinforcement learning agents (e.g., AlphaZero) generate synthetic self-play data to refine policies. "Style" emerges implicitly: fine-tuning on a painter's corpus yields outputs with brush-stroke statistics matching the artist, yet never identical.
Crucially, generation is sampling: given a prompt (conditioning), the model draws from a posterior over latent representations. This process is mathematically a push-forward measure—new instances arise from composing learned functions with noise or interpolations. The question is whether the result ever escapes the convex hull (or manifold) of training experience to constitute invention.
Grounding
Let the empirical training distribution be
$$\mathcal{L}(\theta, \phi) = \mathbb{E}{q\phi(z \mid x)} \left[ \log p_\theta(x \mid z) \right] - D_\text{KL}(q_\phi(z \mid x) || p(z))$$
(the evidence lower bound, ELBO). Sampling
$$\min_{G} \max{D} V(D, G) = \mathbb{E}{x \sim p\text{data}} [\log D(x)] + \mathbb{E}_{z\sim p_z}[\log(1 - D(G(z)))]$$
At equilibrium \(p_\theta \approx p_\text{data}\). Diffusion models formalize creation as score-matching: the reverse SDE denoises from pure Gaussian noise guided by
Under the manifold hypothesis, real data lies on a low-dimensional submanifold \(\mathcal{M} \subset \mathbb{R}^D\) ($D$ high,
Novelty can be quantified rigorously. Define:
$$\text{Novelty}(x) = 1 - \max_i \text{sim}(x, x_i)$$
(where \(\text{sim}\) is cosine similarity in embedding space or inverse Euclidean distance).
Value is domain-specific (e..g, win probability in chess). Creativity score is often the product or harmonic mean; probabilistic models hit an empirical ceiling around 0.25 because extreme novelty correlates with low usefulness. In discrete spaces (chess move sets, token sequences), combinatorial explosion ensures most samples are literally unseen, yet semantically recombined.
Experiments
To ground our discussion in tangible intuition and rigorously test whether generative AI truly creates new knowledge or merely remixes existing patterns, lets do some experiments. These experiments form a deliberate narrative arc that mirrors the actual generative pipeline: we first recover the structure the model learns, then examine how it navigates that structure, then test its capacity for conceptual operations within it, and finally quantify the statistical novelty of the outputs. Each step is motivated by the scientific question raised by the previous result.
Experiment 1: Manifold Learning and Reconstruction
We begin at the foundation. Under the manifold hypothesis, real-world data lies on a low-dimensional submanifold. Does a generative model actually recover this geometry?
To investigate, we synthesize a 3D spiral manifold representing structured "experiences" and trained a VAE to learn its latent representation. After training, we decoded new samples from the latent space. The visualization reveals both the learned manifold surface and where the generated points land.
The result is striking: virtually all generated sampled (red triangles) lie directly on or extremely close to the learned manifold surface (smooth blue curve), while the original training data (blue dots) defines its support. This geometrically confirms our earlier theoretical claim—generative models do not invent new dimensions or escape the support of the training distribution. Instead, they learn a continuous approximation of the data manifold and sample from it. This sets the stage for understanding how novel points are actually produced.
Experiment 2: Latent-Space Interpolation
Given that model learn a manifold, the next logical question is: What mechanism produces outputs that feel new? The dominant process is interpolation.
We selected two distant points on the learned manifold (source and target) and generated a sequence of latent vectors along the linear interpolation path, decoding each into the data space. The animation captures this trajectory in full.
Observe how the dotted interpolation path does not take a straight-line shortcut through the empty space. Instead, it elegantly follows the curvature of the learned manifold. Every intermediate sample is technically "novel"—it was never present in the original training set—yet each remains fully consistent with the geometry the model extracted from data. The animation provides a powerful visual evidence for the claim that most AI generation is sophisticated geodesic traversal rather than genuine invention. It naturally raises the next question: Can models perform more sophisticated operations than simple linear interpolation?
Experiment 3: Vector Arithmetic as Concept Blending
While geometric interpolation explains basic generation, frontier models appear capable of higher-order operations such as analogical reasoning. We tested whether this too can be understood as structured mixing within latent space.
Using a 3D embedding space, we positioned conceptual vectors corresponding to "man", "woman", and "king", then performed the classic vector arithmetic operation (king - man + woman). The animation below traces the complete calculation path and the emergence of "queen" concept.
The resulting point lands on an emergent concept ("queen") that was never explicitly present as a combination in the training data. This is no longer simple spatial interpolation, but semantic blending—a higher-order form of recombination that closely parallels how LLMs and Diffusion models generate novel compositions from learned embeddings. Yet it remains recombination within a fixed learned space.
Experiment-4: Gaussian Mixture Sampling and Novelty Quantification
After observing the mechanisms (manifold learning -> interpolation -> arithmetic), we must answer the crucial quantitative question: How far, in practice, do generated samples depart from the training distribution?
We fit a Gaussian Mixture Model to synthetic training data (two overlapping clusters) and sampled new points from the learned distribution. The left panel shows the original training points versus the generated samples. The right panel displays a histogram of the minimum Euclidean distance from each generated samples to its nearest training neighbor, serving as a direct novelty metric (mean novelty distance marked by dashed line).
The distribution reveals that while some generated points remain close to the training examples, a substantial fraction lies at meaningful distance. This confirms that generative models do produce statistically novel outputs. However, these points remain well within the probabilistic support of the learned distribution—they are novel recombinations, not extrapolations into entirely new regimes.
These four experiments reveal a consistent picture. Generative models first discover the underlying manifold of their training data, then navigate and densify it through interpolation and semantic arithmetic, and finally produce outputs with quantifiable statistical novelty. The entire process is grounded in manifold learning, geodesic traversal, and probabilistic sampling—powerful, elegant and deeply impressive—yet fundamentally rooted in recombination rather than ex-nihilo creation. This progressive visual journey provides the empirical backbone for the conclusion that follows.
Real-World AI Discoveries
Scale and hybrid mechanisms push beyond toy limits. Consider three canonical cases.
AlphaGo's Move 37 (Game 2 vs. Lee Sedol, 2016): Experts assigned this shoulder hit a 1-in-10,000 probability of human play. It emerged from Monte Carlo tree search guided by a policy/value network trained initially on human games then refined via millions of self-play iterations. The move was not in any training corpus; self-play generated new data, and RL exploration discovered a high-value policy deviation. Mathematically, the policy network's softmax over legal moves, combined with MCTS lookahead, sampled an outlier from an evolving distribution—exploratory creativity via synthetic data expansion.
FunSearch (DeepMind, 2023): An LLM generated candidate programs for open mathematical problems (cap sets, bin packing). An evolutionary evaluator selected and mutated survivors, yielding the first LLM-driven scientific discoveries: improved bounds on cap-set size and novel heuristics outperforming human-designed ones. Here, recombination of code fragments (learned syntax and patterns) + selection pressure produced transformational outputs—new algorithms altering the search space itself.
AlphaTensor and AlphaFold: AlphaTensor discovered matrix-multiplication algorithms superior to Strassen's 1969 result after self-play on tensor decompositions. AlphaFold predicted novel protein structures and interactions absent from the PDB, enabling downstream drug design. In both, the models extrapolated within learned physics/chemistry manifolds but reached configurations whose value was unforeseen.
These cases argue that when generation couples with self-play, evolution, or massive scaling, effective novelty emerges. Yet skeptics note: all remain recombinations at root—self-play expands the dataset, not the ontology. The distinction blurs at the frontier scales: if recombination density exceeds human cognitive capacity, functional creation is achieved.
Open Questions
Philosophically, the debate echoes Plato's mimesis versus Aristotle's poiesis, or Kant's genius (originality without rule). Current AI is masterful at mimesis + combinatorics but lacks the "free play of imagination" Kant required for genius. Yet from an instrumentalist view (Dennett), if outputs pass Turing-like creativity tests and advance science/art, the ontological label "mere mix" becomes moot.
Implications abound. In science, AI co-pilots accelerate hypothesis generation (e.g., new materials via graph neural networks). In art, copyright law grapples with derivative status. Epistemologically, AI forces re-examination of human uniqueness: if both species remix, the difference lies in embodiment and intentionality.
Open research directions:
Detect when a model alters its own architecture or discovers new inductive biases.
Benchmark diffusion/VAE models on deliberately held-out manifold perturbations.
Quantify synergy when researchers steer latent traversals.
Extend empirical ceilings beyond 0.25 via world models or meta-RL.
Creation Redefined
Generative AI does not create ex nihilo; it remixes, interpolates, and explores learned manifolds with superhuman efficiency. Experiments and mathematics confirm this. Yet at sufficient scale and with reinforcement or evolutionary wrappers, the outputs achieve functional originality such as new, surprising, and valuable, which match Boden's definition and rivaling human discovery. The distinction between "create" and "mix" dissolves into a continuum. For homo sapiens, the invitation is clear: treat AI not as oracle or parrot, but as a powerful collaborator whose recombination engine can be steered toward genuine frontier expansion. The next move 37 awaits our joint play.



