Into The Infinite: Hilbert Spaces & AI/LLM Search Spaces

An archer smoothly picks up his arrow, bow in hand.
He slots the arrow in place and confidently pulls it back into position.
Taking aim at the target, the archer adjusts his position, the angle of the bow and the tension on the pullback – and silently gauges the wind speed.
In that moment, each one of those variables has a range of values that will be set at the moment of release.
While the real world usually can be approximated by nice round numbers, if you drill down far enough there is virtually an infinite number of possible values those variables can take.
Adjustments of a fraction of an inch or angle (or any other variable) can have an effect on the path the arrow takes on its way toward the target.
Those slight changes in variables may not have a dramatic affect on the eventual placement of the arrow on the target from a macroscopic viewpoint. If you scoped down far enough, each path would have its own unique fingerprint – defined by the accumulated context of variables around the moment of release.
The number of those “fingerprint paths” would approach infinity, with each variable contributing its own set of infinite values to each fingerprint – the path space could take on an infinite dimension of possibilities (more or less).
Spaces with infinite dimensions are difficult to imagine for us spatially (humans are tuned for 3 spatial dimensions plus one time dimension), so studying spaces like this means we have to abstract that spatial understanding.
As it so happens, in the world of mathematics the study of Hilbert Spaces can give us the machinery and intuition necessary to work with very large – sometimes infinite – spaces.
An Informal View Of Hilbert Spaces
If you’re just picking up this blog or want a refresher on some preliminaries, go back and check out my posts on topology, metric spaces, vector spaces and — perhaps most relevant to this post – inner product spaces.
Important for the intuition of Hilbert spaces, inner products, again are a generalization of the familiar dot product – a measure of similarity between two vectors. An inner product space, then, is a vector space equipped with an inner product (that, in turn, has a few important properties you can explore in that previous post).
This allows us to work with concepts like angles, orthogonality, projections, lengths and distances between vectors within those spaces — along with all of the other operations one would expect when working with vectors ( addition, multiplication/scaling, et. al.).
If that space has no “holes” in it – that is – no matter where you find yourself in the space, you’ll have a precise vector to define your location (loosely speaking) – you have what is called a Hilbert Space.
Having no holes in a space that has infinite (or a very high amount of) dimensions means that space is… large (very large).
A (Semi) Formal View Of Hilbert Spaces
The “no holes” characteristic mentioned above is something called “completeness”.
A formal definition of a Hilbert Space is that it is a complete, normed inner product space. Four characteristics essentially — it is a (1) vector space equipped with an (2) inner product and (3) norm (a measure of size of a vector) and that is (4) complete (no holes).
(A side note here that a more abstract space called a Banach Space – one that doesn’t require – loosely – the inner product requirement may be of interest to some folks here.)
Although the theme of the infinite is carrying through this post, Hilbert Spaces can also be finite in nature (not infinite, intuitively).
Applications of Hilbert Spaces can be found everywhere – from signal processing to information and communication theory; quantum mechanics, in particular is the most familiar place you’ll see Hilbert Space applied (for many reasons).
Hilbert Spaces And Graph Theory
The last three posts here on graph theory, eigenvalues and Markov chains would be good to refresh here, but as you’ll see Hilbert Spaces have a natural application to graph theory.
A graph (G) defined by a set of vertices (V) and edges (E) can be represented as a square matrix – with each entry i, j defined as a 1 if vertices i and j are connected – zero if not; something called the adjacency matrix.
Swap the “ones and zeros” with probabilities – the probability of moving from node (vertex) i to node j – you find yourself with a transition matrix.
Expanding the dimension of this matrix up — to infinity — this matrix is constructed and exists in a Hilbert-like space. One that not only tells you what is connected, but also the probability of transitioning (walking) from one particular node in the (very large) graph to the next.
This construction of a transition matrix within a Hilbert-like space contains the entire structure of the graph – which allows mathematicians to analyze and apply different operators (more on this in a future post) to learn more about its underlying characteristics and other hidden structures within that graph.
Connection To SEO & AI/LLM Spaces
Mentioned in the past few posts here, the application of graph theory to SEO is only natural — web pages representing vertices (nodes) and the links between them representing edges.
The adjacency matrix then represents if page connections and the transition matrix tells you the probability of a user moving from one page to the next.
My mental model for AI/LLM responses works similarly, where each token, concept, entity or other particular unit of language represents a walk through the LLMs internals – on a dynamically generated “web graph” of sorts at each step in the generation process as it chooses candidate tokens to fill the response.
These dynamic walks essentially carry their own representative adjacency matrix and transition matrix — ones that are evolving at each time step (expanding and contracting as candidates are added/removed from the available candidate pool according to evolving bases).
Over time, these “walks” have their own transition tendencies (especially through and toward more steady/stable regions – or eigenstates), moving from one token to the next that has a Markovian feel; using prompts and the accumulated context around the prompt (and the previous tokens generated) to produce each state as a response is generated.
Ideally, the space these “walk trajectories” take place in would have Hilbert-like qualities – in particular “no holes” (completeness) – to ensure more reliable/stable responses (in most cases).
While there will always be a response generated (offering a sense of completeness), coherent responses (coherent trajectories) that are relevant to the current context are not always accurate or reliable (thus leaving a “relevance hole” of sorts in the associative space).
Much of our work in SEO is filling or optimizing these relevance gaps (in a sense) – for particular sets of users – and the same thought process can be applied to AI/LLM spaces.
Many more notes on the graph theory & AI/LLM spaces in the coming weeks.
The “Hilbert Archer” & Transition Trajectories
Going back to the archer analogy – I mentioned how even small/fractional adjustments in the archer’s position, bow angle, cord tension and wind speed can yield different arrow paths – each with its own unique fingerprint (when scoped down far enough).
Imagine that same arrow having the ability to pass through multiple targets along a graph of branched targets — and each target has a variable thickness that will divert the arrow along different branches until it hits the end of the target sequence.
That unique arrow path will also tend to pass through some targets more often than others (due to the thickness of each target).
The set of targets are vertices on the graph and the branches of the targets the edges — the arrow represents our “walk” or path through the graph (you can generalize this to a particle for those interested), with thickness representing the transition matrix (raising or lowering the probability of passing to the next target based on where the arrow hits the target).
While this analogy is a bit loose, it should help you visualize “trajectories” rather than “generations” – much like our web surfer passing through web pages.
The uniqueness of each path, the range of context accumulations around the moment of search/prompt – combined with the varying transition tendencies between different contextual states (as each token, word, phrase, et. al. selected) creates the conditions that make measurement quite difficult (and ultimately user-dependent).
Fortunately, with the machinery and intuition gained from Hilbert Spaces (along with some non-classical methods), we should have the tools necessary for not only measurement – but also optimization of – these spaces.
More bits on measurement, the archer and a peek into the looking glass next week.



