Quick answer: Cleaning, formatting and aligning data to a model's documented input requirements using standard techniques, where the transformations are known in advance, is generally unlikely to be a core R&D activity on those facts, subject to the activity's own facts and the statutory tests. A core R&D activity may sit narrower: where it is unknown whether a synthetic-data regime can substitute for scarce real labels, or whether a drift detector can separate genuine shift from sensor fault, determinable only by a systematic progression of work. You self-assess.
18 August 2026 — this article describes the current rules. The 2026-27 Federal Budget announced proposed changes to the R&DTI; the ATO states the measure is not yet law, and industry.gov.au states the changes would apply to income years starting on or after 1 July 2028 if enacted.
Most of the cost of an AI system is spent on its data. Labels are scarce and expensive, the rare class that matters is rare precisely because it matters, and the model that passed acceptance in March behaves differently in September without a line of code changing. Effort is not the statutory test, and the volume of data work in a project says nothing about whether any of it is a core R&D activity.
This article is about the data regime itself — labels, synthetic substitutes and drift. It does not cover running a model on constrained hardware, or the engineering of transport and infrastructure data pipelines; both have their own treatment in our Insights.
What the Official Guidance Already Treats as Ordinary Work
AusIndustry's AI sub-guide lists activities that are generally not core R&D activities. One of them describes the bulk of what most teams call "the data work", verbatim:
"cleaning, formatting, and aligning data to meet a model's documented input requirements using standard data preparation techniques, where the transformations required are known in advance"
Two qualifiers in that sentence carry the weight: standard techniques, and transformations known in advance. Deduplication, schema mapping, unit normalisation, outlier trimming to a documented rule and resampling to a documented cadence are examples where the required transformation may be known in advance. The guide also lists integrating known model outputs such as scores and labels into applications using pre-defined logic, and running regression or acceptance tests using established methods where the expected outcomes are already known.
The same guide is blunt about novelty of tooling: "Using an AI model or technique that is new to you does not, by itself, mean the activity is eligible for the program", and "using AI in software development does not make an activity eligible" (business.gov.au). A first encounter with a synthetic-data library, or with a drift-monitoring package, is a first encounter — not, by itself, an unknown outcome.
Against that, the guide states where AI work may reach the core test: "AI-related activities may meet the requirements of a core R&D activity where a technical hurdle exists and an expert in the field considers that only experimentation will determine if a proposed solution, or the way to develop a solution, can resolve the technical hurdle."
The Core Test, Stated in Full
Core R&D activities are experimental activities whose outcome cannot be known or determined in advance on the basis of current knowledge, information or experience, but can only be determined by applying a systematic progression of work that is based on principles of established science and proceeds from hypothesis to experiment, observation and evaluation, and leads to logical conclusions; and that are conducted for the purpose of generating new knowledge, including new knowledge in the form of new or improved materials, products, devices, processes or services (s 355-25(1), ITAA 1997; business.gov.au).
The established-science limb is the one data work most often cannot reach. The limb does not ask established science to indicate in advance that the hypothesis will hold — if it did, nothing with an unknown outcome could ever qualify; it asks that the progression of work itself be based on principles of established science. A team can run a disciplined loop — generate a batch, retrain, measure, repeat — while what is varied, and what the result is taken to mean, rest on nothing more than trial and error; that loop is a search, not a systematic progression of the kind the section describes. Where a synthetic-data question does rest on established science, it usually rests on something specific: image-formation and radiometric models for rendered visual data, sampling theory and the statistics of covariate shift for tabular data, or the physics of the sensing chain itself.
Where the Data Question Can Become the Experiment
Two situations recur in which the answer is not available from current knowledge, information or experience:
Whether synthetic labels can substitute for scarce real ones: Where the class that matters occurs a handful of times a year, the constraint is not compute or method choice — it is that the real examples do not exist in useful numbers. The literature reports synthetic augmentation results, but on collections, sensing geometries and class definitions that are not yours. The open question is whether a synthetic regime produces samples that are faithful on the dimension the model must actually key on, or samples carrying a generation signature the model learns instead. That question can be posed with a measure and a decision rule, and it can come back "no".
Whether drift can be told apart from a fault in the sensing chain: A monitored input distribution moves. Either the world moved — a new supplier, a new mix, a new season — or a sensor drifted, a firmware update changed a filter, or a camera acquired a smear. The two demand opposite responses, and the observable data can look similar. Implementing thresholds and alerts on a documented statistic using an established package is described by the guidance above as ordinary work. Whether any observable signature separates the two causes on your instrumentation, at a false-alarm rate the operation can absorb, may not be answerable in advance. Retraining on a schedule because performance decayed is not that question.
Exclusions That Sit Across Data Work
Subsection 355-25(2) of the ITAA 1997 lists categories that cannot be core R&D activities at all, whatever they look like. Three sit directly across the analytics roadmap:
Routine testing and analysis: Paragraph (f) excludes activities associated with complying with statutory requirements or standards, including maintaining national standards, calibrating secondary standards, and routine testing and analysis of materials, components, products, processes, soils, atmospheres and other things — which may cover sensor calibration or routine production testing where those activities are associated with complying with a statutory requirement or standard.
Reproduction of a commercial product or process: Paragraph (g) excludes any activity related to the reproduction of a commercial product or process by a physical examination of an existing system, or from plans, blueprints, detailed specifications or publicly available information. Both elements have to be present: the thing reproduced must be a commercial product or process, and it must be reproduced by one of those means. Reimplementing a published synthetic-data or drift-detection method from its paper and reference code plainly goes to the second element; whether the first is met depends on whether what is being reproduced is a commercial product or process on the facts, which a published method with reference code is not automatically.
Internal-administration software: Paragraph (h) excludes developing, modifying or customising computer software for the dominant purpose of internal administration of the developer's business functions, or those of an entity connected with it or an affiliate.
See what does not qualify. Annotation campaigns, extract pipelines and evaluation harnesses may qualify as supporting activities where they are directly related to a core R&D activity. Supporting R&D activities are activities directly related to core R&D activities; but where an activity is of a kind referred to in s 355-25(2), or produces goods or services, or is directly related to producing goods or services, it is a supporting R&D activity only if undertaken for the dominant purpose of supporting core R&D activities (business.gov.au). Where a labelling pipeline produces, or is directly related to producing, goods or services, the additional dominant-purpose test applies if it is being assessed as supporting R&D.
A Worked Hypothetical: Can Synthetic Defects Substitute for 190 Real Labels?
Hypothetical and illustrative. The figures are invented to show the shape of an experiment; nothing here indicates that any activity is eligible.
An Adelaide packaging manufacturer inspects foil seals inline. The defect that matters — a micro-crease that passes visual inspection but fails under pressure — appears about once in 12,000 units. Eighteen months of production across three lines yielded 190 confirmed labelled examples.
Baseline: A classifier trained on the 190 real positives reaches recall 0.62 at precision 0.90 on a sealed holdout of 640 images containing 58 real defects.
Requirement, recorded 9 February: A synthetic regime is accepted only if it reaches recall ≥ 0.85 at precision ≥ 0.90 on that sealed holdout and holds recall ≥ 0.80 on line 3, whose illumination geometry is deliberately excluded from the generator's calibration. If it misses either, the regime is rejected.
The search, recorded the same week: Published augmentation and sim-to-real results for rare-class visual defects; the two vendor toolkits already licensed; the internal capture logs. None reported results for a specular, semi-transparent crimped foil under strobe illumination, where the highlight pattern — not the crease — dominates the image. Hypothesis: crease defects rendered from a measured surface reflectance model, under each line's measured strobe geometry, are faithful enough on the crease dimension to substitute for real positives. Basis: radiometric image formation and the statistics of covariate shift.
The trials:
Classical augmentation: (elastic warps, photometric jitter, cut-and-paste). Recall 0.71 at precision 0.90 on the validation split. Better than baseline, short of target.
Generative model: (fine-tuned on 190 real crops). Failed. Recall 0.83, but line 3 recall 0.44 (baseline 0.58). Probe classifier separated synthetic from real at 0.97 AUC. Generator reproduced the illumination signature, and detector keyed on it.
Physics-based rendering: (parametric crease geometry on measured reflectance model). Recall 0.79; line 3 recall 0.76; synthetic-versus-real separability 0.61 AUC. Transfer restored, target still missed.
Rendered positives plus domain randomisation: Validation recall 0.88. Sealed holdout: recall 0.87 at precision 0.91, and line 3 recall 0.85. Both targets met.
Result reached: A process window exists, but it is bounded: within the range tested, the regime held both requirements up to a target point; a generatively synthesised regime did not, for an identified reason.
Activity boundary: The experimental work runs from the dated requirement and search to the one read of the sealed holdout. The rendering harness, the reflectance measurement rig and the annotation of the 190 real defects are ancillary work built so the trials could be run and scored. Rolling the accepted model into the line controller, the operator interface, scheduled retraining, and routine testing of production units afterwards sit outside it. Whether any of this is registered, and on what basis, is for the company to self-assess with its own advisers.
Where an RSP Fits
AusIndustry describes Research Service Providers as scientific or technical service providers a company can engage to conduct R&D activities on its behalf, registered in specific fields of research (business.gov.au). There is also a threshold point for smaller claimants: Qualifying expenditure incurred to a non-associate RSP may still form part of the offset where total notional deductions are below the usual $20,000 threshold, provided the services relate to a research field for which the RSP is registered. Using an RSP does not guarantee eligibility — you still self-assess. An RSP supplies research capability, not tax advice.
See claiming R&D under $20,000. Offset rates, the refundable and non-refundable tiers and the entitlement rules are covered separately in refundable vs non-refundable offset.
Frequently Asked Questions
Q: Is data cleaning and preparation eligible for the R&D Tax Incentive?
A: Generally not as a core R&D activity. AusIndustry's AI guidance lists "cleaning, formatting, and aligning data to meet a model's documented input requirements using standard data preparation techniques, where the transformations required are known in advance" among activities that are generally not core R&D activities. Such work may be a supporting R&D activity where it is directly related to a core activity and, where the dominant-purpose trigger applies, undertaken for the dominant purpose of supporting one. You self-assess.
Q: Is generating synthetic training data an eligible R&D activity in Australia?
A: It depends on what is unknown. Running an established generator to a documented recipe is generally unlikely to be a core R&D activity on those facts, subject to the activity's own facts and the statutory tests. A core R&D activity may exist where whether a synthetic regime can substitute for scarce real labels at a stated operating point could not be determined in advance on current knowledge, information or experience, and could only be determined by a systematic progression of work based on principles of established science, conducted for the purpose of generating new knowledge.
Q: Is monitoring a model for drift claimable under the R&DTI?
A: Implementing logging, alerts, dashboards or routine performance checks using established methods and tools to confirm expected model behaviour is listed in the official guidance as generally not a core R&D activity. The position can differ where the unknown is whether any observable signature distinguishes genuine distribution shift from a fault in the sensing chain, at a false-alarm rate the operation can absorb — but that question must be investigated through the required systematic progression of work.
Q: Is a data labelling pipeline core or supporting R&D?
A: Usually neither automatically. Annotation and pipeline work is normally considered against the supporting-activity provision: it must be directly related to a core R&D activity, and where it produces goods or services, is directly related to producing goods or services, or is of a kind referred to in s 355-25(2), it qualifies only where undertaken for the dominant purpose of supporting core R&D activities. Where a pipeline produces, or is directly related to producing, goods or services, the dominant-purpose test also applies if the pipeline is being assessed as supporting R&D.
Sources & Further Reading
legislation.gov.au — Income Tax Assessment Act 1997 — Div 355, incl. ss 355-25 and 355-30
Related: R&D for software and AI · what does not qualify · what an RSP is · claiming R&D under $20,000 · refundable vs non-refundable offset · Insights
Talk to Ignition Research before you register or lodge. As a Registered Research Service Provider at Lot Fourteen in Adelaide, we assist transport, utility and asset-analytics teams to separate the platform build from the experimental question, set the baseline, measure and hold-out protocol in advance, and record the trial sequence as it happens. Your company self-assesses and remains responsible for its own claim; we are not a registered tax agent. See also refundable vs non-refundable offset and R&D Tax Incentive in Adelaide. Get in touch.
This article is general information from a Registered Research Service Provider about the R&D Tax Incentive. It is not tax, legal or financial advice; eligibility depends on your circumstances and you should self-assess and seek your own advice.
Thinking about a project like this?
If you're weighing up an AI, software or technical improvement project and can't tell yet whether it's implementation or research, start with a quick read on where it sits.

