Module 2 — Data Privacy, Confidentiality, and Professional Responsibility
The module that governs all the others
The workflows in the second half of this course will save you real time. None of them are worth using if they expose client data, breach your professional obligations, or put a filing in the hands of an unreviewed text generator. This module sets the guardrails that make everything else safe. Treat it as the operating envelope: any use of AI in your practice must stay inside these boundaries, and where a later workflow seems to push against them, the boundary wins.
Two forces make this non-negotiable for accountants specifically. First, we hold data that is among the most sensitive a person or business owns — Social Security numbers, full financial pictures, tax return data, health-related figures, the intimate details of a company's operations. Second, we are members of a licensed profession with codified duties of confidentiality and due care. A marketing team experimenting with AI risks an embarrassing draft. An accounting firm experimenting carelessly with AI risks a confidentiality breach, a professional-standards violation, and the trust that is the entire basis of the client relationship. The stakes are simply different, and the discipline has to match.
What confidentiality means when the data leaves your machine
Start with a mental model of where your words go when you use a cloud AI tool. You type into a chat box in your browser. That text travels over the internet to the provider's servers, where the model runs and generates a reply that travels back. Your input does not stay on your computer. It is transmitted to, and processed by, a third party. Everything about responsible AI use for accountants follows from taking that sentence seriously.
Once data has left your machine and reached a provider, several questions determine whether you have a problem: Is the transmission encrypted? Does the provider store your input, and for how long? Can the provider's employees see it? And — the question that surprises people most — will your input be used to train future versions of the model? The answers are not the same across tools, and critically, they are not the same across tiers of the same tool. This is the distinction that most firm accidents turn on, so we treat it carefully.
Consumer tiers versus enterprise tiers
Most popular AI assistants come in more than one tier, and the tiers differ in ways that have nothing to do with how good the answers are and everything to do with what happens to your data.
A consumer tier — the free version, or a low-cost individual personal plan — is built for individuals experimenting with the technology. Its terms of service frequently permit the provider to retain your conversations and to use your inputs to improve and train their models. That last point deserves to land fully: on a consumer tier, the client trial balance you paste in to get a variance narrative may become training material, which means fragments of it could influence the model's future outputs to other users, entirely outside your control. Even where training is off by default or can be toggled, consumer tiers commonly retain conversation history on the provider's servers, accessible to the provider, for some retention period. Consumer tiers are also where "human review for quality" clauses live — provisions letting provider staff read a sample of conversations. None of this is nefarious; it is ordinary product development. It is also flatly incompatible with client confidentiality.
An enterprise or business tier — sold to organizations, typically under a negotiated or standardized business agreement — is built for exactly this problem. The defining commitments of a reputable enterprise tier are that your inputs are not used to train the provider's models, that data retention is limited and often configurable (including zero-retention or short-retention options), and that the arrangement is governed by a business contract with data-protection terms rather than a consumer clickthrough. Many enterprise offerings will also sign a data processing agreement and, where relevant, a business associate agreement, and provide administrative controls, access logging, and single sign-on. The practical upshot: enterprise tiers can be configured so that client data is processed without being retained for training and without lingering indefinitely, which is the baseline you need before any client data touches a tool.
The rule that follows is blunt and worth memorizing:
Never place client-identifying information, return data, or other protected data into a consumer or free AI tool. Client data belongs only in an enterprise-tier tool that your firm has vetted, configured to disable training on your inputs, and covered by an appropriate business agreement.
Do not assume a tool's tier from its name or its price. Read the terms for the specific tier you are on, find the data-retention and model-training provisions, and check the account's settings — many tools expose a toggle for "improve the model for everyone" or "training," and on any tool touching client work that toggle must be off, backed by contract language that makes the setting meaningful.
Retention and training settings, concretely
Because settings are where good intentions succeed or fail, here is what to actually verify before a tool is cleared for client-adjacent work, framed as a checklist you can run:
- Model-training on your inputs is disabled, and disabled by the account terms, not merely by a personal preference toggle that could reset. On enterprise tiers this is typically the contractual default.
- Data retention is understood and acceptable. Know whether inputs are stored, where, and for how long. Prefer options offering short or zero retention for sensitive processing. "We don't train on it" and "we don't keep it" are two different promises; get both.
- Human review of conversations is understood. Know whether provider staff may view content and under what controls. Enterprise agreements typically constrain this sharply.
- Access within your firm is controlled. If the tool stores conversation history, know who in the firm can see whose conversations, and whether that matches your internal confidentiality partitions.
- The provider's security posture is documented. Encryption in transit and at rest, and a recognized security attestation, are table stakes for a vendor you will route client-adjacent data through.
Run this before adoption, and re-run it when a provider changes terms, which they do. A tool that was safe last year under one contract is not automatically safe this year under an updated one.
Minimize even inside a safe tool: de-identification and data minimization
Choosing an enterprise tier is necessary but not sufficient. The most robust habit, and the one that protects you even when a tool turns out to be less safe than believed, is to minimize the sensitive data you send in the first place. Ask, for every prompt: does the model actually need this specific identifier to do the task? Usually it does not.
A model drafting a variance narrative needs the numbers and the account structure; it does not need the client's name, the taxpayer identification number, or the specific address. A model triaging an asset list into recovery-period buckets needs the asset descriptions and amounts; it does not need the owner's name or Social Security number. You can strip or mask identifiers before sending and re-insert them yourself afterward. Replace "Riverside Dental Group, EIN 12-3456789" with "the client" or a neutral label; replace real names with placeholders; remove account numbers and personal identifiers that play no role in the transformation you are asking for. This practice, de-identification paired with data minimization, means that even in the worst case — a misconfigured setting, a provider breach, an unexpected retention — what escaped was not identifiable to a person or a return. It is defense in depth, and it is cheap: usually a few seconds of find-and-replace, or a habit of describing rather than naming.
Be realistic about what de-identification can and cannot do. Stripping a name helps; but a sufficiently detailed financial picture can sometimes be re-identified from its particulars, and true return data — the actual figures on a filing — carries obligations regardless of whether a name is attached. So minimization is a strong supplement to using a properly configured enterprise tool, not a substitute for it. The order of operations is: use a vetted enterprise tool, and minimize what you put into it.
Subprocessors, data location, and the questions behind the questions
Two further considerations separate a superficial tool review from a rigorous one, and both surprise firms that thought "enterprise tier" was the end of the analysis.
The first is subprocessors. When you send data to an AI provider, that provider frequently relies on other companies to deliver the service — cloud hosting, infrastructure, sometimes the underlying model itself if the product is a wrapper around another company's model. Each of those is a subprocessor that may touch your data, and a rigorous review asks who they are and what commitments flow down to them. Reputable enterprise agreements disclose subprocessors and bind them to compatible data-protection terms; a consumer tool typically discloses little and binds nothing you can rely on. This is why the "friendly accounting label" caution from later in the course matters: a specialized accounting AI product may route your data through a general model provider as a subprocessor, and the terms that govern that leg of the trip are the ones that actually protect your client. Read for the subprocessor chain, not just the front door.
The second is data location and cross-border transfer. Where, physically and jurisdictionally, is your client's data processed and stored? For many engagements this is a compliance question with real weight — certain data carries residency or transfer constraints, and moving it to a provider's servers in another jurisdiction can implicate obligations beyond your professional standards. You do not need to become a data-protection lawyer, but you do need to know enough to spot when an engagement's data has location sensitivity and to route it accordingly, using tools that let you control or confirm the processing region, or keeping it local. When in doubt on a sensitive engagement, the local-model option — where data never leaves your hardware at all — sidesteps the entire cross-border question, which is one more reason it belongs in a firm's toolkit.
Behind both considerations sits a single principle worth stating directly: you remain responsible for client data even after it leaves your hands. Handing information to a vendor does not hand off the obligation. The vendor's commitments are the mechanism by which you keep your obligation, which is exactly why the terms, the subprocessors, and the data location are your business and not merely the vendor's. A breach at a provider you chose carelessly is, in the ways that matter to a client and to your standards, your breach.
Professional responsibility: the standards framing
Confidentiality is not merely good practice for accountants; it is a professional duty. The AICPA Code of Professional Conduct establishes a member's obligation to maintain the confidentiality of client information, and separate rules govern the disclosure and use of client data, including specific obligations around taxpayer information. You do not need to memorize section numbers to act correctly; you need to internalize that routing client information to a third-party AI provider is a disclosure of that information to a third party, and that such disclosures must be consistent with your confidentiality obligations, your engagement terms, and any consents you have obtained. An enterprise tool with proper data-protection terms is how you keep that disclosure consistent with your obligations; a consumer tool with training enabled is how you make it a violation.
Two professional duties beyond confidentiality bear directly on AI use. The first is due care and competence — the obligation to perform services with the diligence and skill the work requires. Using a tool you do not understand, or accepting its output without the review that its known unreliability demands, is inconsistent with due care. Competence here includes understanding the tool's failure modes, which is precisely why Module 1 spent its length on hallucination. The second is the broader duty to serve the client's interest and the public trust, which is undercut the moment an unreviewed, possibly-fabricated AI output reaches a client as if it were your professional work.
There is also a client-facing dimension. Whether and how your firm uses AI in performing an engagement is something clients may reasonably expect to be addressed. This is where engagement letters and firm policy come in.
Engagement letters and firm policy
The clean way to handle the disclosure question is to address it in your engagement letter and your firm policy, rather than deciding it ad hoc at the keyboard. An engagement letter can disclose, in plain terms, that the firm may use technology tools, including AI-assisted software, in performing the services, that such tools are used under the firm's supervision and review, and that client confidential information handled through them is protected by appropriate safeguards and agreements. It can also, where the firm chooses, obtain the client's acknowledgment or consent to the use of such tools. Coordinating your engagement language with counsel is prudent; the point for this course is that the disclosure belongs in a document the client agreed to, not in a decision an individual makes silently on a Tuesday afternoon.
A firm AI-use policy turns the principles in this module into rules people can follow without re-deriving them. A workable policy answers, at minimum: which tools are approved and on which tiers; what categories of data may and may not be entered, and into which tools; the requirement that a qualified person review every AI-assisted output before it is used or delivered; the de-identification expectation; who is accountable for vetting a new tool before anyone adopts it; and how the firm handles a suspected exposure. We build a concrete version of this checklist in Module 6. Introduce it here as the mechanism that scales confidentiality from your own careful habits to a whole staff, because the least experienced person on your team, on their busiest day, is where a data-handling policy is actually tested.
The human-in-the-loop principle, made specific
Module 1 introduced the principle that a qualified professional reviews and takes responsibility for every AI output. In the language of professional responsibility, this is not optional and not delegable to the tool. State it in full because the workflows ahead depend on it:
The AI assists; the CPA is responsible. The model produces a draft, a summary, a triage, a proposed classification. A qualified professional then reviews that output against the source material and against their own knowledge, corrects what is wrong, supplies what is missing, and makes the professional judgment the output feeds into. The professional's name and license stand behind the result, exactly as they would if a junior staff member had prepared the draft. The AI is, in this framing, a very fast junior preparer with an unpredictable error pattern and no accountability of its own — which is exactly why the review cannot be skipped and cannot be a rubber stamp. A reviewer who signs off on AI output without checking it against the source has not reviewed anything; they have laundered a text generator's guess into a professional deliverable, and they own the consequence.
This is why every workflow in Modules 4 and 5 ends with a review-and-guardrail step, and why we treat that step as the most important part of the workflow rather than an afterthought. The prompt is the easy part. The review is the profession.
Where this leaves the cost segregation example
Return to the anchor. When AI triages an asset list into proposed depreciation buckets, several of this module's rules converge on it at once. The asset descriptions and amounts can go to a vetted enterprise tool; the property owner's name, entity identifiers, and any personal identifiers should be stripped first, because the bucketing does not need them. The output — proposed five, seven, fifteen, and long-life classifications — is preliminary and gets reviewed by a qualified professional, never filed as generated. And the study that will actually support a filing is a full engineering-based cost segregation study performed by qualified professionals; the AI's triage is a preliminary aid that speeds the front of that process, not a replacement for it. Every guardrail in this module is present in that one small workflow, which is why it makes such a good teacher.
Recap
Client data leaving your machine reaches a third-party provider, and confidentiality depends entirely on what that provider does with it. Consumer and free tiers frequently retain inputs and may use them to train models; they are unsafe for client data. Enterprise and business tiers can be configured not to train on your inputs, to limit retention, and to operate under a business agreement — that configuration is the minimum bar before client data touches a tool. Verify training and retention settings explicitly, and re-verify when terms change. Minimize and de-identify what you send even inside a safe tool, as defense in depth. Professional standards — confidentiality, due care and competence, and the duty to the client and public — govern all of this; routing client information to an AI provider is a disclosure that must be consistent with your obligations. Address AI use in engagement letters and codify it in a firm policy. Above all, keep the human in the loop: the AI assists, the qualified professional reviews every output and remains responsible for it. With these guardrails fixed, the next module turns to the practical skill that makes AI genuinely useful — prompting.