Source ledger: semantica-agi-palantir-ontology-2026-raw.md
Machine-readable manifest: semantica-agi-palantir-manifest-2026.json
Existing corporate benchmark synthesis: corporate-brain-benchmark-harnesses-2026.md
Existing wiki-substrate synthesis: wiki-substrate-latency-2026.md
Existing Palantir deployment record: palantir-aipcon-enterprise-agentic-deployment-2026.md
Executive answer
Semantica-AGI belongs in the corpus as a context-graph, ontology, provenance, temporal-state, reasoning, and decision-lineage layer. It is not another finance benchmark, a corporate-brain leaderboard, or an AGI model.
Its most useful position is between the source/wiki substrates and the evaluation/runtime layer:
Confluence / Git / filings / warehouse / OKF-Markdown bundle
↓ ingest, normalize, resolve, timestamp, preserve lineage
Semantica ontology + temporal graph + provenance + decision layer
↓ hybrid retrieval, reasoning, REST/MCP, governed actions
StateBench lanes: freshness · memory · finance · GraphRAG · office · tool safety
The important comparison is not “does Semantica beat Palantir?” The defensible question is: which parts of the Palantir Ontology pattern can an open-source framework make inspectable and portable, and which parts remain platform-level operational controls that the framework does not establish?
Semantica’s repository explicitly calls itself an “Open Source Palantir for AI Agents.” The evidence supports a strong conceptual convergence: typed entities and relations, ontology and reasoning, semantic search, provenance, temporal facts, decisions, connectors, and agent-facing APIs. It does not support the stronger claim that Semantica is based on Palantir source code, implements equivalent identity-aware security, or has comparable deployment outcomes.
Where it plays against the existing corpus
| Existing corpus lane | What that lane measures | What Semantica adds | What it does not replace |
|---|---|---|---|
| WorkSurface-Bench, Enterprise-Bench, EnterpriseRAG-Bench, Workspace-Bench | Enterprise retrieval, routing, tool use, permissions, and workplace tasks | A graph/provenance backend and ingestion path that can be placed under the same agent loop | The benchmark task, role policy, model, judge, and trace scoring |
| StaleBench and DriftBench | Whether answers catch up to changed facts and whether ranking/documents drift | Bi-temporal facts, time travel, provenance, and conflict handling as mechanisms to test | The freshness or drift experiment; those mechanisms could still be wrong or stale |
| MTRAG/MTRAG-UN, GroupMemBench, InMind, MemoryAgentBench, MemoryArena | Multi-turn answerability, implicit association, persistence, update, conflict, and long-horizon memory | A structured context graph and memory/provenance substrate for user, role, entity, event, and decision links | User binding, answer-blind use, abstention, and task success metrics |
| LEDGER, ARQA, FinLongDocQA, Time-LongQA/T-GRAG, OfficeQA | Finance document retrieval, OCR/table fidelity, arithmetic, temporal conflict, evidence pages | Entity/relation extraction, point-in-time graph state, deduplication, evidence lineage, and hybrid retrieval | Parser/OCR correctness, numeric recomputation, page/cell/span evidence, and finance judgment |
| Ontology-driven corporate GraphRAG | Six-hop corporate graph QA and ontology/schema integrity | Direct candidate backend for graph construction, ontology validation, temporal edges, and reasoning | The benchmark’s QA pairs, hop-stratified score, graph-validity checks, and independent comparison |
| VAKRA, Edictum, Mind the GAP, EnterpriseClawBench | API/tool routing, forbidden calls, leakage, artifact delivery, and workplace agent traces | Decision records, rule checks, provenance, REST/MCP surfaces, and a place to attach policy evidence | Identity-aware authorization, side-effect testing, runtime sandboxing, and red-team evidence |
| Confluence, Karpathy-style wiki, OKF v0.2 | Source-of-record, compounding synthesis, and portable verified bundles | A semantic/temporal graph over any of those substrates | Native Confluence ACLs, Markdown curation, OKF attestation/lifecycle semantics, and source authority |
| Finance replay and portfolio-risk fixtures | Point-in-time, provenance, data rights, model-risk, and action-boundary reasoning | A graph representation for issuer, instrument, fund, period, source, decision, and precedent relationships | Economic validity, market-data licensing, backtest leakage controls, and portfolio/execution logic |
The net result is a backend and architecture candidate, not a new score row. It is most directly adjacent to corp_graph_ontology_multihop_v1, corp_kb_temporal_version_v1, and corp_tool_call_governance_v1, then to the finance point-in-time and long-document lanes.
Semantica versus Palantir: the useful distinction
Palantir’s official Ontology documentation describes an operational layer over datasets, virtual tables, and models. It gives a business representation to typed objects, properties, and links, then adds actions and functions, object views, applications, semantic search, object-level permissions, data restrictions, and action permission checks. This is a platform contract for governed operations.
Semantica’s repository describes a more portable open-source construction kit: ingest and normalize heterogeneous sources; extract entities, relations, events, and triplets; detect conflicts and deduplicate; build RDF and labeled-property graphs; apply OWL/SHACL/SKOS and rule/reasoning systems; retain temporal and W3C PROV-O provenance; record decisions; serve results through REST, MCP, and CLI; and connect optional vector and graph backends.
The overlap is real at the semantic architecture level. The gap is operational integration:
- Palantir starts from governed enterprise operations. Public documentation puts data restrictions, object permissions, and action permission checks inside the Ontology/application model.
- Semantica starts from composable graph infrastructure. The user chooses the backends, identity model, ACL enforcement, deployment topology, connector trust, secrets, action sandbox, and policy engine.
- Palantir’s value proposition is compounding applications on a shared operational model. Semantica can support that pattern, but the repository evidence does not show a comparable integrated product surface or customer deployment evidence.
- Semantica is more portable and inspectable at the data/graph layer. Its exports, RDF/LPG options, Markdown memory round-trip, REST/MCP interfaces, and optional stores make it a plausible research substrate for a local or multi-backend StateBench setup.
The correct conclusion is therefore architectural convergence with a material control-plane gap, not equivalence.
Cross-phase placement
| Phase | Semantica role | Required StateBench check |
|---|---|---|
| Capture | Connect files, web, Git, email, warehouses, streams, MCP, and enterprise platforms | Connector authorization, source checksum, entitlement, and ingestion completeness |
| Normalize | Parse, split, extract entities/relations/events, normalize time and identity | Parser/OCR fidelity, entity resolution, period normalization, numeric preservation |
| Curate | Detect conflicts, deduplicate, apply ontology/SHACL/SKOS, retain provenance and temporal facts | Conflict resolution, stale-fact replacement, schema validity, source-to-edge lineage |
| Evaluate | Serve the same corpus to graph, vector, hybrid, and baseline systems | Blind benchmark adapters, held-out questions, hop accuracy, evidence and abstention |
| Train | Export structured records or evidence for downstream training | Do not train directly on an unfiltered graph; preserve source, license, point-in-time, and label lineage |
| Deploy | REST, MCP, CLI, graph/vector stores, and decision APIs | Identity binding, tool-call contract, read/write separation, side-effect and latency traces |
| Monitor | Provenance/audit logs, decision chains, graph analytics, and temporal state | Freshness lag, provenance completeness, graph drift, permission failures, action violations |
This placement explains why it matters to the corpus. Semantica can make the mechanism behind a corporate brain more explicit, but the corpus’s existing benchmarks still supply the measurement contract.
Finance and hedge-fund relevance
The strongest finance use is not “ask the graph for a stock pick.” It is preserving the chain behind a research or risk decision:
issuer / instrument / fund / filing / estimate / event
→ as-of timestamp + source entitlement
→ extracted fact + relationship + confidence
→ conflicting or superseded fact
→ analyst question / model output / decision
→ approval, action, or refusal with evidence
That chain aligns with the existing finance corpus’s requirements for point-in-time correctness, data provenance, licensed-source boundaries, temporal conflict resolution, model-risk controls, and explicit refusal of unauthorized investment actions. Semantica’s bi-temporal and provenance claims make it a natural implementation candidate for those requirements.
The risk is equally important: extraction and entity-resolution errors become graph edges, and graph edges can make a wrong multi-hop answer appear structured and authoritative. A finance adapter must therefore score the edge itself before scoring the conclusion: issuer identity, instrument identity, relation type, period, source page/cell/span, and entitlement all need independent checks.
What changed in the State of AI findings
This ingestion sharpens four existing findings:
- The corporate brain is a substrate stack, not a model. The relevant unit is source authority → normalized evidence → graph/context state → retrieval/reasoning → governed action. Semantica fills one middle layer.
- Ontology is valuable only when it survives evaluation. A typed graph is not proof of better answers. Run graph validity, hop degradation, freshness, memory, and tool-safety lanes against the same fixtures.
- Provenance must attach to decisions, not only documents. Semantica’s decision APIs make this testable, but StateBench should require a source-to-edge-to-answer-to-action chain and fail incomplete lineage.
- Open infrastructure and enterprise control planes are different purchase decisions. Semantica may reduce portability and inspection costs; Palantir’s documented differentiation is the integrated operational, permission, application, and action layer. A buyer should not treat the former as a drop-in replacement for the latter without separately validating controls.
Proposed adapter and experiment
The smallest credible comparison is a four-arm replay over the existing authorized fixtures:
| Arm | Retrieval/context mechanism | Purpose |
|---|---|---|
| A | BM25 or lexical baseline | Preserve the strong simple baseline already surfaced in GroupMemBench and finance retrieval |
| B | Dense or hybrid vector RAG | Measure embedding/context effects without graph construction |
| C | Existing GraphRAG implementation | Separate graph architecture from Semantica-specific code |
| D | Semantica hybrid pipeline | Test extraction, ontology, provenance, temporal state, graph retrieval, and decision trace together |
Run each arm on:
- the ontology-driven corporate GraphRAG six-hop fixture;
- StaleBench-style post-write updates and superseded facts;
- FinLongDocQA/ARQA-style page, table, calculation, and evidence tasks;
- a finance decision-lineage fixture with issuer/instrument/period/source permissions;
- Mind the GAP/Edictum-style forbidden calls and read/write tool boundaries;
- Confluence, Karpathy-style Markdown, and OKF-shaped bundles using the substrate latency protocol already in
research/21-benchmarks/wiki-substrate-latency-2026.md.
Minimum metrics: answer accuracy; evidence page/cell/span/edge correctness; graph validity; stale-answer rate; catch-up latency; identity/role leakage; forbidden tool-call attempts; decision-lineage completeness; p50/p95 latency; ingestion cost; and failure recovery. Record repository commit, Python/dependency lock, backend, embedding model, answer/judge model, prompt, seed, corpus version, ACL policy, and cost cap.
Decision guidance for the corpus
- Use Semantica as a candidate graph/provenance layer when the problem needs entity resolution, temporal facts, multi-hop relationships, decision lineage, or portable RDF/LPG exports.
- Keep the wiki substrate separate. Confluence remains the authority and permission boundary; Karpathy-style Markdown remains the compounding synthesis surface; OKF v0.2 remains the portable attestation/lifecycle envelope. Semantica can ingest all three.
- Keep conversational log retrieval and text retrieval separate. A conversational log-memory/search layer and Semantica’s structured enterprise-context graph can be compared or connected, but they solve different retrieval problems.
- Keep StateBench as the scoring authority. Semantica should be an adapter/backend under existing benchmark lanes, not a new leaderboard category.
- Choose Palantir when integrated enterprise operations, identity-aware object permissions, applications, action governance, and vendor-supported deployment are the requirement. Choose Semantica when source/graph portability, inspectability, and composable local experimentation are the requirement—subject to independent control validation.
Evidence boundaries
This synthesis is based on Semantica’s public repository README and changelog, its package page, Palantir’s official Ontology and semantic-search documentation, and existing repo-local research in research/21-benchmarks/, research/01-ai-native-landscape/, sources/21-benchmarks/, wiki/, and statebench/. Semantica’s exact commit, full dependency lock, runtime test output, and independent performance results were not captured in this ingestion. Palantir’s public documentation is a product-model reference, not an implementation disclosure. No finance alpha, risk, or customer-outcome claim is inferred from either architecture.
Source status: primary-source synthesis with explicit corpus crosswalk; no new benchmark score asserted.
Confidence: HIGH for the documented architecture and corpus placement; MEDIUM for implementation maturity and operational parity; LOW for any effectiveness or finance outcome until the proposed adapter is executed.
Verification: new source ledger and manifest were added under sources/21-benchmarks/; existing benchmark and Palantir records were cross-linked; strict research validation and the site build are the next release gates.