The long-form pieces. How we think about getting AI to work in places where the answer matters and the rules live in people’s heads.
When an expert overrides an AI system, that correction is not noise. It is a teacher. We study how to turn corrections into rules the system can apply next time, without retraining.
A model is only as good as what you put in front of it. We study how to pick the right context: the right documents, the right examples, the right structure. Accuracy comes from what the model sees, not how big it is.
Operators rarely write down what they do. We study how to watch real work and turn it into rules an AI system can follow. No interviews, no workshops, no policy documents.
Fine-tuning is expensive and hard to undo. We study how a living knowledge graph can turn a general-purpose model into a domain expert at the moment it answers, with no model weights touched.
We benchmarked Phyvant's knowledge layer against five graph-RAG and agent-memory systems on identical tasks and the same model backbone. Phyvant scored 0.957 to the field's 0.31. On the custodial work that decides whether enterprise AI is auditable at scale, it wasn't close.
Every tweak to a production AI pipeline once meant paying to re-run the model on thousands of examples. Here is how we cut that to a fixed one-time pilot, so teams ship a fix the same day instead of waiting two weeks on a GPU bill.
Your AI retrieval dashboard reads healthy while one slice of users gets wrong answers 7× more often than the rest. Here is why averaging hides the failure, what it costs in production, and how we close the gap.
We collaborate with researchers and teams building serious enterprise AI.
Get in touch