arXiv:2609.00581nvidia/nemotron-3.5-lightning-30b-a3bSeptember 1, 2026

Enoki: Efficient Multi-Level Hallucination DetectionExplained for Beginners

Elisei Rykov, Timur Ionov, Nikolay Ivanov +5 more

Computation and Language

Abstract

Ensuring factuality remains a critical challenge for deploying LLMs in high-stakes settings. Existing hallucination detectors usually operate at a single level: claim-level methods provide interpretable factual units, while span-level methods localize unsupported text. Bridging these views is costly, as LLM-heavy pipelines require multiple decomposition and verification calls, and modular systems need additional claim-to-span alignment. We propose Enoki, an Open Information Extraction framework for multi-level hallucination detection. Enoki extracts text-anchored relational facts, verifies them against evidence, and projects unsupported facts back to hallucinated spans. This shared representation enables claim-level verification and span-level localization without requiring separate alignment. Enoki supports LLM-based, encoder-based, and rule-based extraction regimes, balancing accuracy and inference cost through a common interface. Experiments show that Enoki remains competitive with strong claim-level systems while using fewer resources and achieves superior performance on fine-grained span- and entity-level localization. We also release EnokiQA, a dual-granularity dataset with aligned claim-level verification and span-level localization annotations.

Here is a structured explanation of the paper Enoki: Efficient Multi-Level Hallucination Detection.

1. The Problem

The core challenge this paper addresses is the deployment of Large Language Models (LLMs) in high-stakes environments—such as medical diagnosis, legal analysis, or research assistance—where "hallucinations" (fluent but unsupported statements) can have serious consequences.

Current hallucination detectors usually operate at only one level of granularity, creating a gap between interpretability and precision:

  • Claim-level methods decompose an answer into factual units (claims) and verify each one. This is great for providing interpretable evidence, but it relies heavily on the quality of the decomposition. If a claim is poorly constructed (e.g., missing a crucial modifier or temporal context), the verifier might judge a different proposition than what the user actually wrote.
  • Span-level methods localize the exact text fragment responsible for the unsupported content. This is useful for editors who need to fix the text, but span labels alone don’t explain why the text is wrong or what factual structure was violated.

Bridging these two views has been costly. Previous approaches required separate pipelines: first decompose the text, verify the claims, and then perform a second, expensive step to align those claims back to the original spans. This modular approach is fragile, often propagating errors from the decomposition stage into the alignment stage.

2. How It Works (The Technical Mechanics)

Enoki proposes a unified framework based on Open Information Extraction (OpenIE). Instead of treating a sentence as a single block or a bag of claims, Enoki extracts text-anchored relational facts—structured triples of (Subject, Predicate, Object) that remain tied to specific locations in the source text.

Here is the mechanics broken down into a concrete analogy for a software engineer or product manager:

The Analogy: The "Fact Card" System Imagine you are writing a report and you want to fact-check it. Instead of just highlighting sentences (span-level) or writing high-level summaries (claim-level), Enoki works like this:

  1. Fact Extraction (The "Card" Creation): Enoki reads each sentence and pops out "Fact Cards." Each card contains a relational triple, such as Enoki | is | mushroom or Tesnière | was born | in Montpellier. Crucially, these cards are "text-anchored," meaning they carry the exact character positions of the subject and object within the original sentence.
  2. Verification (The "Check"): These fact cards are then checked against a reference context (like a Wikipedia article or a retrieved document). The system asks: "Does the evidence support the claim that 'Enoki is a mushroom'?" It uses a natural language inference (NLI) verifier to score this.
  3. Projection (The "Link"): If a fact is deemed unsupported (a hallucination), Enoki doesn't just flag the whole sentence. Because the fact card is anchored to the text, it can project the "unsupported" part back to the exact span. In the analogy, if the verifier says "Enoki is a mushroom" is unsupported, Enoki highlights the specific text span "Enoki" or the incremental addition that caused the error.

Why this matters: Because the facts remain tied to the text throughout the process, the system achieves claim-level verification (we know which facts are true) and span-level localization (we know exactly where the error is) from the same intermediate representation. There is no need for a separate, error-prone "alignment" step.

Enoki offers three "engines" for this extraction, balancing accuracy and speed:

  • ENOKI-LLM: Uses a powerful LLM to extract facts. Highest accuracy, but higher cost.
  • ENOKI-ENCODER: A trainable encoder-based model (using ModernBERT). A middle ground that is significantly faster than LLMs.
  • ENOKI-RULE: A deterministic, rule-based system. Fastest and cheapest, though slightly less flexible than learning-based methods.

3. Key Results & Benchmarks

The experiments demonstrate that Enoki achieves a strong balance between performance and efficiency.

Entity-Level Localization (HalluEntity):

  • ENOKI-LLM (with incremental prompting) achieved an AUPRC of 79.70, significantly outperforming the best implicit verifier (LLaMA3.1-8B-Instruct at 72.03) and explicit OpenIE baselines (which hovered around 58-67).
  • Even the rule-based and encoder-based variants remained competitive, showing that the "fact representation" is the key driver of performance, not just the LLM backend.

Span-Level Localization (MuSHROOM, RAGTruth, PsiloQA):

  • The paper uses Span Coverage F1, a metric that rewards finding unsupported content even if the system's prediction is more "fine-grained" than the (often coarse) gold annotation.
  • ENOKI-LLM performed best on MuSHROOM and PsiloQA, indicating that high-capacity decomposition helps recover fine-grained hallucination arguments.
  • ENOKI-ENCODER and ENOKI-RULE remained competitive, achieving much of the LLM's benefit at two orders of magnitude lower latency.

Computational Efficiency:

  • On the RAGTruth benchmark, ENOKI-ENCODER achieved 69.1% F1 in just 0.13 seconds.
  • It was 4–10× faster than competitive baselines and up to two orders of magnitude faster than multi-stage LLM pipelines.
  • The efficiency gain comes from the fact that verification is done with a lightweight encoder model (ModernBERT-large) rather than autoregressive generation through an 8B+ parameter LLM.

4. Why It Matters (Key Takeaways)

  • Unified Representation: Enoki solves the "granularity gap." By using text-anchored facts as the shared currency, it provides both interpretable verification and precise localization without the overhead of modular alignment. This makes the system more robust and easier to maintain.
  • Flexible Accuracy-Efficiency Trade-off: The release of three backends (LLM, Encoder, Rule) means practitioners can choose the right tool for their budget. If you have a tight budget but need better performance than a rule-based system, ENOKI-ENCODER offers a compelling middle ground, preserving most of the localization benefits of the LLM version.
  • The ENOKIQA Dataset: The paper introduces a new benchmark (3,990 labeled examples) with aligned claim- and span-level annotations and long-form answers. This fills a critical gap in the field, providing researchers with a resource to train and evaluate multi-granular hallucination detectors.
  • Limitations to Watch: The system's performance hinges on the fact extraction stage. If the extractor misses a key proposition or merges distinct facts, the verifier cannot recover it. Additionally, because the current implementation primarily works sentence-by-sentence, it may struggle with cross-sentence phenomena like coreference (pronouns referring to previous mentions) without the optional decontextualization step.

In summary, Enoki represents a shift toward multi-granular hallucination detection that is both more precise (localizing errors to specific spans) and more practical (offering efficient, modular backends), making it a significant step toward deploying reliable LLMs in real-world applications.

Want to understand AI papers like this from scratch?

Follow the free AI Learning Roadmap →