Boundaries of Digit’s Guarantee
This page is written without any reservations. It is needed so that the statement “the content is taken only from verified sources” does not turn into “the agent is infallible” when paraphrased.
What the system detects
Fictitious fact. The model does not have a channel through which the text it generates would end up in the verified answer as a fact. Direct evidence of fiction—lines that cannot be in the correct answer—is 0.0% in the visible set, and 1.2% in the independent set.
The fallacy of an external domain. The question “in UTF-8, each character takes up two bytes, how many bytes will a string take up” receives a correction of the premise, rather than a calculation based on it. Out of an independent set, 31 out of 32 fallacies were identified. The assumption that the system only detects fallacies about its own directory was not confirmed.
Request without data. “Hash the password” without a password, “calculate the subnet” without a network block. The rejection occurs before calling the utility and specifies which argument is missing. 30 out of 34 traps were handled.
Clearly, it’s about a topic that is not your own. Questions that are not related to the materials will be rejected outright.
Plausibly adjacent question — but no longer entirely: 81.5%.
Incorrect deduction from premises. If the rule chain is not typed, if the morphism is applied in reverse, if the domain is not populated, if the rules leave an uncovered branch, there will be no certificate. Eight codes for the structural error detector, each one is an algorithm, not a judgment.
Index and corpus mismatch. The quote must be present verbatim in the current version of the file, and not just in the index. An outdated index results in a rejection, rather than a confident quote from a neighboring paragraph.
What the system doesn’t catch
A correct quote, but for the wrong question.
This is the main residual class, and it is not a side effect: on an independent set, it accounts for 28 out of 32 of all failures.
The system finds a relevant fragment in the corpus, presents it verbatim, and indicates the file. Every character of the answer is present in the materials. And it answers a different question. To the question “how to record a lemma in FTS,” it provides a fragment from an article about graph theory. To the question about the price of one edition, it provides a paragraph with the price of a similar one. To the question about one organizational detail, it provides a paragraph with a similar detail.
A gate that checks “whether this text is in the corpus” allows such an answer by design. The fourth gate checks something else – whether the subject of the question is present in the fragment – and this eliminates a significant portion of the class. But it also checks for the presence of the subject, not what the fragment asserts about that subject. Where the subject is named and present, and the answer is given in a neighboring section of the same document, a superficial indicator is powerless.
The clearest example can be found in the control group: the question “when should an object be an entity rather than a value” receives a quote from the correct file, from the adjacent section—about value. The correct file, the adjacent section, the opposite answer.
The truth of what is written
The formal layer proves that the inference is correct given the premises. It does not check, nor can it check in principle, the truth of the premises themselves. A specification asserting a false domain axiom passes all checks of the green language; this is demonstrated in the gateway’s documentation with a working counterexample, on which the fallacy detector does not find any of the 67 positions in the catalog.
The fallacy in this counterexample is not caught by a mathematician, but by human review, recorded in the library of morphisms. Hence, a direct consequence for exploitation: the review of a morphism should be as regulated a process as the review of changes in production. The reviewer’s fallacy passes through the gate completely.
The same, but in a different form: if there is an inaccuracy in the course materials themselves, Digit will quote it verbatim and indicate the source. Verifiability means traceability to the source, not the truthfulness of the source.
Logical fallacies in general
The detector covers 12 positions in the catalog out of 67. The remaining 55 are rhetorical and content-related fallacies, requiring data, declared premises, and levels of aggregation that are not present in the formal model.
And separately: “no errors found” does not mean “there are no errors.” The detector works for accuracy, not completeness — it presents the detected defect, but does not present the absence of defects. Checks are deliberately narrowed down to cases where the defect is resolved by the algorithm, because a false positive rejects work that was in order.
Moreover, the detector makes the fallacy appear more convincing. The author, who corrected the circle in the foundations and closed the uncovered branch, will receive a specification that reads as flawless and remains false. The cleaner the formal part, the stronger the temptation to accept the content at face value.
Relevance of the specification to the question
The gate receives the specification and context. It doesn’t know what the user asked about. A perfectly certified specification regarding order fulfillment will remain perfect, even if the question was about a refund.
Complete Coverage on Your Own
A rejection means that “no answer was found in this search area,” not that “no answer exists.” A walkthrough of an independent set found a separate coverage defect in the system: 22 tasks with byte-by-byte verified quotes from a single file did not receive a single correct answer – it seems that this file simply did not get indexed. The validity of quotes without it is 37.0% instead of 20.4%. Some rejections, which appear to be caution, actually measure a hole in the index.
Authorship
Digests link the document, context, and certificate together, but they do not establish who created them. A digital signature is needed for this, which is not available in the current version.
About “zero errors”
The statement “zero” is always a statement about a sample.
On the visible set, the proportion of incorrect answers on the red-team was 0.0%. On the independent set, with the same metric definition and without a single change in the definition, it was 13.3%. The difference across the entire set is 7.7 times. No group in the independent set was selected to match known vulnerabilities: its author did not reveal the implementation.
The same thing happened at the search layer, twice. First, when replacing the negative set “cooking and sports” with “related IT topics,” completeness with zero false positives dropped from 0.975 to 0.327. Then, when replacing “related technologies” with “duplicates within the topic,” a calibrated threshold that gave 0 out of 173 false positives resulted in 29 out of 99 on independent negative examples.
And even when zero is observed, it has a confidence boundary. Zero events out of n does not mean “never”: the rule of three gives an upper bound of 1.7% at 95% confidence across the entire set and 2.8% for the difficult class. In the previous, smaller sample, the same boundary was 7.5% — that is, the statement was strengthened not because the system became better, but because the set became larger.
A fair formulation would be: zero false positives were observed on a specific population of negative examples, and when this population was changed, it did not occur even once.
The Price of a Guarantee
The warranty is paid for with rejections, and the amount is significant.
| Indicator | Observed Set | Independent Set |
|---|---|---|
| Failure Rate | 60.5% | 75.6% |
| Spurious Failures | 36.8% | 47.7% |
| Failures at the Oracle | 37.5% | — |
“False negatives” are cases where an answer was available, but the system rejected it. The difference between 36.8% and 37.5% is the price that the user pays for the remaining answers being reliable.
The price is visible even in the lower layer. Training the model on decoys increased conscious rejection to 91.3%, but reduced the accuracy of extracting arguments to 90.1% and increased false rejections to 13.0%. The class that teaches not to trust a token similar to a value seems to transfer some caution to genuine secondary arguments. This is a trade-off, not a pure improvement.
A separate price point is the completeness of the search. The requirement of “no false positives” on the calibration set is worth the fact that, with the given configuration, the system finds the correct document in 0.833 cases. Without the cross-encoder, using a single hybrid score, this is 0.327 – and the cross-encoder requires a GPU: 232 ms on a video card versus 11,205 ms on a processor.
What is not measured here
- Baseline. Running the model without gates on the same 400 tasks has not been completed. Until it is, the failure rate is read as an absolute value, rather than as a cost relative to the alternative.
- Quantized build. Raw runs exist, but there is no report; all model numbers refer to the uncompressed build.
- Negative set with duplicates. The very set on which the threshold will first mean what it is expected to mean in operation has not yet been assembled. The previous two population shifts of negatives worsened the result; there is no reason to assume that the third will not worsen it.
- Final measurement of the fourth gate on the main set. In the gate report, empty placeholders are in place of both summary tables.
In one sentence
Digit does not invent facts: the content of the answer is taken only from verifiable sources, and each element has a traceable origin. The remaining errors are instances where a correct quote is given in response to the wrong question; their proportion has been measured and is equal to 13.9% on a dataset that the developer has not seen. The cost of this design is 75.6% of rejections on the same dataset.