Quick answer: “The model got better” is not, by itself, enough to evidence a systematic experiment. Restate the unknown as something measurable on a defined document population, then consider building an adjudicated reference set labelled independently by two or more qualified reviewers and measuring how much they agree. That study quantifies ambiguity and label noise in the reference data, which helps make later scores interpretable. It does not, by itself, prove statutory technical uncertainty — that is the separate question of whether a competent professional could have determined the outcome in advance from existing knowledge. Pre-specify and time-stamp the threshold, report intervals alongside point estimates, hold out a test set and log failed configurations.
As at 27 July 2026, the Australian Government had announced reforms to the R&D Tax Incentive in the 2026–27 Federal Budget, intended to apply to income years starting on or after 1 July 2028. Until those changes take effect, the program continues to operate under the current legislation. See the proposed $50,000 Budget measure.
Legal-tech teams get stuck at the same place. The eligibility test for a core R&D activity wants a hypothesis, an experiment, an observation, an evaluation and a logical conclusion — the sequence business.gov.au sets out on its conducting core R&D activities page. Fine, if you are measuring tensile strength. But if you are building clause extraction, matter triage or a drafting assistant, what you are trying to predict is a senior practitioner's judgement, and many legal-judgement tasks do not have a single uncontested reference answer. There is no thermometer for "is this indemnity clause unusual enough to escalate?" So the honest problem is not "how do I write this up later" but what would a real experiment even look like here?
We are Ignition Research, a Registered Research Service Provider (RSP000047) at Lot Fourteen in Adelaide. We work on the science side — activities, evaluation protocols, evidence trails — not tax: your claim position belongs to your own registered tax agent, and you self-assess. Whether the AI work is eligible at all is a prior question, covered in is the AI work eligible and, for this vertical, in legal tech.
Step 1: Restate the Unknown as Something a Number Can Move
"Our model gives better legal answers" is not a hypothesis: no dependent variable, no population, no comparison. business.gov.au says a hypothesis should explain what result you aim to achieve, how you intend to achieve it, and why that result may or may not be achievable, developed before you start, and is blunt that selecting from available options by trial and error does not meet the requirement.
Method note: The controls, holdout sets, pre-specified thresholds and measurement plans described below are research and evidence practices. They are not all express statutory requirements. The statutory eligibility tests remain those in Division 355.
Turning a judgement task into a measurable one takes four decisions, in writing, before any run:
The population: Which documents — Australian commercial leases over a defined value band? A sampling frame, not "our data".
The label schema: A closed set a qualified reviewer can apply consistently — escalate / note / no action per clause, plus the character span it attaches to. Vague labels produce vague results and, as below, reviewer disagreement that means nothing.
The metric: Agreement with the adjudicated reference labels; class-level recall where a miss is expensive; span-overlap for extraction. Pick them before you see results, and decide up front how you will express uncertainty.
The comparator: A baseline you already have — the current rules engine, an off-the-shelf model, a junior reviewer's pass. AusIndustry's AI-related activities guidance does exactly this: repeated runs comparing a new method against a baseline on the same data, assessed for statistical significance.
That guidance also lists adjusting parameters using established methods, where the effect of the change is well understood, as unlikely to meet the core R&D requirements. Prompt-tweaking until the demo feels better is not an experiment; testing a proposition about why a defined method should resolve a defined technical hurdle is.
Step 2: Build an Adjudicated Reference Set
If your reference labels come from one lawyer, they reflect that lawyer’s answers rather than an independently validated reference set. You also cannot readily distinguish model error from variability in the labelling process.
Say what the artefact is. It is an adjudicated reference set: labels produced by a defined process, on which a defined group of qualified people converged. It is not objective truth, and calling it a "gold standard" smuggles in a claim you cannot support. Every score is a score against that reference set, and inherits its assumptions.
1. Sample properly
Stratify so rare-but-important categories appear.
2. Label independently and blind
Two or more qualified reviewers, same schema, same written guidance, no visibility of each other's work or of the system output. Record who was trained on what, and when — reviewer training is a variable, not a constant.
3. Measure inter-reviewer agreement before you measure the model
Report raw agreement and a chance-corrected statistic, because on a skewed label distribution raw agreement flatters.
4. Adjudicate
A third qualified reviewer resolves the disputes to produce the reference labels — and keep the disagreement log, an observation in itself.
Pick the right agreement statistic: Cohen's kappa is the one everyone names, and it is correct only for two raters on unordered categories. Match the statistic to the design:
Cohen's kappa: two reviewers applying nominal, unordered labels to the same items.
Weighted kappa: two reviewers applying ordered labels, where different degrees of disagreement should receive different weights.
Fleiss' kappa: a multi-rater measure for nominal labels, commonly used where more than two reviewers rate each item.
Krippendorff's alpha: a flexible option for different numbers of reviewers, missing ratings and different levels of measurement, provided an appropriate disagreement function is specified.
Report whichever you use with an interval — a bootstrap CI resampled over documents is the practical approach — and name the statistic and its weighting in the protocol. A kappa quoted bare to two decimals implies precision a few hundred clauses cannot support.
Low agreement is information about the reference set, not proof of technical uncertainty. It should prompt a review of the schema, reviewer guidance, training and document context before the model results are interpreted.
The Distinction That Matters Most: Two Different Questions
“Two qualified lawyers could not agree on the right answer. Therefore the outcome could not have been determined in advance by a competent professional. Therefore there is technical uncertainty.”
That does not follow. Reviewer disagreement can arise from an unclear label schema, a legal question that genuinely admits more than one professional judgement, inconsistent reviewer training, incomplete document context, or plain ambiguity in the source drafting — and only sometimes from a technical problem the system cannot yet solve. Only that last case is anywhere near what the statute asks about, and disagreement alone does not establish even that. A schema so vague that two reviewers read it differently produces spectacular disagreement and no technical uncertainty whatsoever.
Keep the two questions separate everywhere in your records:
Key distinction: Inter-reviewer disagreement quantifies ambiguity and label noise in the reference set. It helps define a credible evaluation benchmark, but it does not by itself establish the statutory technical uncertainty. The separate R&DTI question is whether a competent professional could have known or determined the technical outcome of the proposed method from existing knowledge without experimentation.
That second question is the one s 355-25(1) of the Income Tax Assessment Act 1997 poses: activities whose outcome cannot be known or determined in advance on the basis of current knowledge, information or experience. business.gov.au glosses current knowledge on its core-activities page as knowledge publicly available or reasonably accessible anywhere in the world at the time you start.
The evidence for it is a different pile of paper from the agreement study: a dated search of published literature, vendor documentation and available benchmarks showing what was and was not already solved, plus a statement of the specific technical hurdle — the retrieval strategy, the representation, the decomposition — and why existing approaches do not resolve it. Mind the subject matter too: s 355-25(2)(d) excludes research in social sciences, arts or humanities. Assess and describe the activity according to its actual substance; a computational or information-retrieval experiment is different from legal scholarship, but its classification is not determined by wording alone.
The agreement study still earns its place. It gives you a defensible measurement instrument, stops you over-reading a score, and evidences the observation-and-evaluation limb of the systematic progression. It is strong evidence about the method — not proof of eligibility, and no single artefact is: you self-assess against Division 355, and nothing here guarantees an outcome. Nor is the agreement figure a ceiling; it is a reference band, and a system scored against adjudicated labels can legitimately score higher.
Step 3: Pre-Specify and Time-Stamp the Threshold, Then Hold Out a Test Set
Write down, dated, before the trial runs: the hypothesis in the what/how/why form above; the metrics and the acceptance threshold; how you will express uncertainty — interval method, resampling unit, and the minimum gap you will treat as meaningful; the parameters you will vary, hold constant and measure, the framing business.gov.au uses for what experiment records should explain; and how often you will evaluate against the held-out set, and who signs off.
Pre-specify and time-stamp means dating the protocol somewhere it cannot be silently backdated — a commit, a document system with version history, a signed file. It does not mean lodging anything with a public registry; there is no registry for this.
Two practical notes for document AI: Split by matter or document, never by clause — clauses from the same agreement leak, and a leaked split produces a beautiful number that means nothing. And lock the test set: repeated peeking compromises its independence and may make it part of the development process. A threshold set after reviewing the results is exploratory rather than genuinely pre-specified.
A Worked Example
Illustrative only. Every figure below is hypothetical — not a projection, and not a statement that any activity is eligible. You self-assess and confirm your position with your registered tax agent.
"Torrens Clause", an Adelaide legal-tech company, is building non-standard-clause detection for commercial leases. Existing tooling handles named clause types but fails on the materially unusual drafting the firm's partners escalate on sight.
Unknown (the statutory question): whether a retrieval-conditioned comparison against a drafted-norms corpus can identify materially unusual clauses in Australian commercial leases, where published approaches and vendor tooling — surveyed and dated before the work started — resolve only named clause types, and no accessible source indicates whether the approach transfers to unnamed unusual drafting.
Reference set: 400 clauses, stratified across six clause families. Two senior property lawyers label independently against a written schema. Raw agreement 84% (bootstrap 95% CI approximately 80–88%, resampled by lease rather than clause, since clauses within a lease are not independent); An unweighted Cohen’s kappa of 0.61 is reported, although weighted kappa would be more appropriate because the label schema is ordered. A third partner adjudicates the 64 disputes.
What that 84% is and is not: It measures ambiguity and label noise in the reference set and sets the band against which later scores are read. It is not evidence that a competent professional could not have determined the outcome in advance — that is the dated survey above.
Threshold, pre-specified and time-stamped: supported if held-out agreement reaches 80% or more with escalate-class recall at or above 0.90, both with 95% bootstrap intervals.
Runs: five configurations over two months, split by matter. Configurations 1–4 fail — two below the recall floor, one collapsing on a clause family, one beating the baseline but not significantly. Configuration 5 reaches 79% agreement on a 120-clause holdout (95% CI approximately 72–86%) and 0.91 recall on 44 escalate instances (95% interval approximately 79–96%). Each run is logged with date, settings, result and the decision it drove.
Conclusion: not accepted on the pre-specified terms. The 79% point estimate did not meet the 80% acceptance rule, while the 72–86% interval shows that the study was too imprecise to conclude that the underlying agreement was below 80%. The trial was underpowered for the threshold chosen, itself a recorded result that can inform the next design.
The failed configurations form part of the systematic progression and should be recorded. business.gov.au states plainly that R&D activities may still be eligible even where you do not reach a positive outcome, and the software development sector guide asks for records of the actions taken to overcome the technical hurdle including attempts that failed and why those actions were unsuccessful. Whether the offset ultimately claimed is refundable or non-refundable turns on aggregated turnover and control: see refundable vs non-refundable offset.
One caution: A qualified reviewer applying a defined label schema is measurement; a user telling you whether they would buy it is market feedback, which business.gov.au's excluded activities page puts under the s 355-25(2) market research exclusion. See what does not qualify.
Where an RSP Fits
A research service provider may assist with designing the evaluation protocol, agreement study and dated knowledge search. Building these elements into the work and recording them contemporaneously generally provides stronger evidence than reconstructing them at year end. Records must be kept for 5 years after you claim the expenditure, per business.gov.au's record keeping guidance.
A threshold point matters for early-stage legal-tech, where first-year spend is often small. Total notional R&D deductions for the income year must generally be at least $20,000 (business.gov.au), and RSP-conducted eligible R&D activities can be claimed even where the usual $20,000 R&D expenditure threshold is not met, per business.gov.au's guidance on getting help from a research service provider. Read that precisely. Where total notional deductions are below A$20,000, the offset base is generally limited to qualifying expenditure incurred to a non-associate RSP for services within a field for which it is registered, together with eligible CRC Program contributions. Other in-house amounts do not automatically form part of that below-threshold offset base. We unpack it in claiming under $20,000 with an RSP. And to be clear about the limit of that: using an RSP does not guarantee eligibility — you still self-assess.
The 2026-27 Budget announced proposed R&DTI changes for income years starting on or after 1 July 2028. They are not current law; see our dedicated Budget update for the proposed measures and their status.
Frequently Asked Questions
Q: Does reviewer disagreement prove technical uncertainty for the R&DTI?
A: No. Inter-reviewer disagreement quantifies ambiguity and label noise in the reference set — it can come from an unclear schema, uneven reviewer training, missing document context, ambiguous drafting, or a legal question that honestly admits more than one professional judgement. It helps you build a credible benchmark, but it does not by itself establish the statutory technical uncertainty. That separate question, under s 355-25(1) ITAA 1997, is whether a competent professional could have known or determined the technical outcome of the proposed method from existing knowledge without experimentation — evidenced by a dated search of what was already publicly available or reasonably accessible. You self-assess.
Q: How do you write a testable hypothesis for an AI document-review project?
A: State what result you aim to achieve, how, and why it may or may not be achievable — bound to a defined document population, a closed label schema, named metrics and a named baseline, written and time-stamped before the work starts. "The model will improve" is not testable; "method X will reach agreement within N points of the reviewer agreement band on population P" is.
Q: How do you measure R&D results when the correct output is a human judgement?
A: Have two or more qualified reviewers label a stratified sample independently and blind, then measure agreement between them with raw agreement plus a chance-corrected statistic matched to the design — Cohen's kappa for two raters on unordered labels, weighted kappa where labels are ordered, Fleiss' kappa for more than two raters, Krippendorff's alpha where there is missing data or mixed label types. Adjudicate disputes to form the reference labels, and report every figure with a confidence or bootstrap interval so a one-point gap is not mistaken for a result.
Q: Is user testing with clients an eligible R&D activity?
A: Testing whether clients like a product goes to consumer interest and preferences, which business.gov.au lists under the market research, market testing and market development exclusion in s 355-25(2), so it cannot be a core R&D activity. It may qualify as a supporting R&D activity if directly related to a core activity and for the dominant purpose of supporting it. Keep it separate from the evaluation protocol.
Sources & Further Reading
Conducting core R&D activities — business.gov.au
Software development sector guide — business.gov.au
AI-related activities and the R&D Tax Incentive — business.gov.au
Excluded R&D activities — business.gov.au
Record keeping — business.gov.au
Check if you are eligible — business.gov.au
Get help from a research service provider — business.gov.au
Income Tax Assessment Act 1997, s 355-25 — legislation.gov.au
Related: legal tech · is the AI work eligible · what does not qualify · under $20,000 with an RSP · refundable vs non-refundable · Budget update
Talk to Ignition Research before you register your activities or design the trial. Evaluation protocols and knowledge searches are generally easier to substantiate when completed and recorded contemporaneously. Get in touch.
This article is general information from a Registered Research Service Provider about the R&D Tax Incentive. It is not tax, legal or financial advice; eligibility depends on your circumstances and you should self-assess and seek your own advice.
Thinking about a project like this?
If you're weighing up an AI, software or technical improvement project and can't tell yet whether it's implementation or research, start with a quick read on where it sits.

