Blog
UnlikelyAI

Introduction
Large language models are remarkably good at producing fluent text. They can summarise documents, draft code, and hold conversations that feel coherent and informed. But despite these surface-level capabilities, LLMs do not understand the world in any structured sense.
Language models have become so good at sounding intelligent that we’ve collectively forgotten to ask whether they’re actually thinking at all.
This isn’t just a philosophical quibble, it’s a fundamental limitation that shows up whenever you need AI to do more than generate plausible text. Ask an LLM to maintain consistency across a conversation, apply logical rules systematically, or explain its reasoning in a way you can verify, and you’ll quickly find yourself in a hall of probabilistic mirrors where nothing is quite as solid as it seems.
What is Symbolic AI?
Symbolic AI takes the approach that dominated the field for decades before neural networks swept. Symbolic systems represent knowledge as explicit relationships and reason through formal logic. They’re transparent, consistent, and verifiable. They’re also brittle, labour-intensive to build, and terrible at handling the ambiguity that makes human language so expressive.
So here we are: LLMs offer flexibility, scale, and linguistic competence, but lack guarantees; symbolic systems offer precision and reliability, but struggle with the ambiguity and richness of natural language.
At UnlikelyAI, we’re working on what might be the most consequential, and most difficult problem in AI today: actually combining these approaches in a way that gives us the strengths of both without inheriting their weaknesses.
Why building Neurosymbolic AI is harder than it sounds?
The challenge isn’t just technical complexity, the real issue is that LLMs and symbolic AI operate on fundamentally incompatible principles:
LLMs are probabilistic. They predict the next word based on patterns in training data, building representations that are distributed across millions of parameters. Knowledge isn’t stored anywhere specific; it emerges from the statistical structure of the model. Their knowledge is implicit, distributed across billions of parameters, and encoded as statistical regularities rather than explicit facts or rules. As a result:
• They generate responses by likelihood, not logical necessity.
• They do not maintain a stable internal model of the world.
• They can produce inconsistent or contradictory answers to the same question.
• They struggle to justify why a particular conclusion follows from given premises.
Symbolic AI treats language as logic. Knowledge exists as explicit facts and rules. This means:
• Conclusions are logically grounded.
• Inferences are reproducible and consistent.
• Reasoning steps can be inspected, audited, and explained.
• Errors are localised and diagnosable.
One system thinks in probabilities and patterns; the other thinks in symbols and logic. The gap between is beyond a missing technical feature, it’s a conceptual chasm.
The translation problem
If you want to combine these approaches, you need translation. Natural language has to become symbolic representation, and symbolic conclusions have to become natural language again. This sounds straightforward until you realize that translation between these two domains is doing most of the conceptual heavy lifting.
UL serves as a translation layer. Natural language inputs are mapped into formal, structured representations that preserve meaning while enabling logical manipulation. These representations can be stored, queried, and reasoned over using symbolic methods, ensuring consistency and correctness.
In this architecture:
• LLMs are used where they excel: interpreting natural language, handling ambiguity, and generating fluent responses.
• Symbolic systems are used where failure is unacceptable: reasoning, validation, constraint enforcement, and explanation.
• Outputs are not just generated, but derived through explicit reasoning steps.
Crucially, UL is not a prompt or a post-processing trick. It is a formal language designed to make meaning computable. This allows AI systems to reason about what has been said, not just echo patterns from training data.
The technical challenges of integration
The technical challenges go deeper than the translation problem:
Combining LLMs with symbolic reasoning introduces several deep technical challenges:
Generative vs. Structured Cognition
LLMs operate in a continuous, probabilistic space. Symbolic reasoning operates in discrete, rule-based systems. Translating between these modes without loss of meaning or correctness is non-trivial.
Consistency Guarantees
LLMs are inherently stochastic. Symbolic systems demand determinism. Ensuring that probabilistic language understanding feeds into logically consistent reasoning requires careful architectural boundaries.
Semantic Fidelity
Natural language is ambiguous, context-dependent, and underspecified. Mapping it into formal representations without stripping away nuance or introducing hidden assumptions is one of the hardest problems in AI semantics.
Solving these challenges is not about maximising benchmark scores. It is about building systems that behave predictably under pressure, where errors are costly and trust is non-negotiable.
Why this work matters
The AI industry has an unhealthy obsession with benchmark performance. We celebrate systems that score 90% on some dataset without asking whether that 90% represents genuine capability or just expensive curve-fitting. The problem with focusing on benchmarks is that they encourage solutions that work for the benchmark rather than for the underlying challenge.
Combining LLMs with symbolic AI isn’t about hitting a score on some test. It’s about building systems that can:
• Reason transparently – not just generate plausible explanations, but show actual logical derivations that can be audited
• Maintain consistency – give the same answer to the same question, not because of memorisation but because of structured knowledge
• Learn systematically – acquire new knowledge in ways that integrate with existing understanding rather than just getting statistically associated with it
• Operate reliably in high-stakes domains – healthcare, law, finance, critical infrastructure—places where “the model said so” isn’t good enough
This matters because the next generation of AI applications won’t be chatbots and content generators. They’ll be systems that need to make verifiable decisions based on complex reasoning over structured knowledge. The current generation of LLMs, for all their impressive capabilities, simply can’t do this reliably.
The real challenge: making structure flexible
Here’s what I think many people miss about this problem: the difficulty isn’t just making LLMs more structured or making symbolic AI more flexible. The difficulty is that increased structure usually means decreased flexibility, and vice versa. You’re not trying to move along a single axis; you’re trying to expand into a space that seems to have contradictory requirements.
Nature solved this problem through a kind of layered architecture human cognition operates with both fast, intuitive pattern matching (analogous to what LLMs do) and slow, deliberate logical reasoning (analogous to symbolic systems). But we don’t yet understand how these layers interact well enough to simply copy the design.
What we’re building at Unlikely AI is an attempt to find a computational version of this layered architecture not by mimicking human neurobiology but by finding formal principles that achieve similar functional properties. Universal Language is our bet that you can construct a representation that’s structured enough for formal reasoning yet flexible enough to capture natural language semantics.
Whether we’re right is an open question. But if we’re not, someone else will need to solve this problem, because the current trajectory of just making language models bigger isn’t going to get us there.
Conclusion
The future of AI shouldn’t be measured by how many tokens a model can process or how many billions of parameters it has. It should be measured by whether these systems can actually think—not in some mystical sense, but in the practical sense of applying logical reasoning to structured knowledge while still understanding the nuance and flexibility of human language.
We’ve spent the last decade proving that statistical learning can scale to extraordinary levels. Now we need to prove that we can combine that learning with the kind of structured reasoning that actually makes systems reliable, explainable, and trustworthy.
This is hard. Harder than most people realise. But it’s also necessary. The alternative is a future where AI systems are powerful but fundamentally opaque, capable but not trustworthy, fluent but not reliable, impressive but not dependable.
At Unlikely AI, we’re building something different. Not because it’s easy, but because the problem is worth solving.
Want to see how we’re tackling one of AI’s hardest problems? Follow our Lab newsletter and learn more about UnlikelyAI’s approach to combining reasoning with language understanding: lab.unlikely.ai




