Module 6 — Building an AI-Enabled Firm Workflow, Safely and Cheaply
From individual skill to firm capability
The previous modules gave you a personal skill: you can now use AI well and safely at your own desk. This module is about the harder thing — turning that into a firm capability that is cost-controlled, policy-governed, adopted by staff who did not take this course, and improving over time. That is a change-management problem as much as a technology one, and the firms that get value from AI are the ones that treat it that way rather than buying a subscription and hoping.
We will cover tool selection on a budget, the cost intuition that keeps spending trivial, a firm AI-use policy checklist you can adopt, the change-management realities, how to measure return honestly, and a maturity path from first experiment to embedded capability. Throughout, the two constraints from earlier modules remain the floor: confidentiality is non-negotiable, and a qualified professional reviews every output.
Tool selection on a budget
You do not need to spend much to get most of the value, and you should resist the instinct to buy the most expensive option for every use. Think in three tiers matched to the work.
A vetted enterprise chat assistant is the foundation. One reputable enterprise-tier general assistant, configured per Module 2 — training on your inputs disabled, retention understood and limited, a business agreement in place — covers the large majority of the workflows in this course. This is the tool client-adjacent work flows through, and paying for the enterprise tier here is not optional spend; it is the price of doing this safely at all. Budget for it deliberately and standardize on it so you are vetting and governing one tool well rather than three tools poorly.
Cheaper and smaller models handle the well-structured, low-ambiguity tasks — categorization against a fixed list, reformatting, simple extraction, meeting-notes transformation. Many providers expose these at a fraction of the cost of their frontier models, and routing easy work to them keeps spending negligible. The skill is recognizing which tasks are "easy" in this sense: they have a right answer, a clear format, and low stakes, and Modules 4 and 5 flagged them as you went.
Local or on-premise models deserve real consideration for a firm, because they change the confidentiality calculus. A model that runs on your own hardware sends nothing to any provider — the data never leaves your control — which makes it attractive for sensitive transformation work where you would otherwise agonize over a cloud tool's terms. Local models have historically lagged frontier models in raw capability, but for the well-structured tasks above they are frequently good enough, and the privacy benefit is substantial. A pragmatic firm posture: local models for routine, sensitive, well-defined transformation; a vetted enterprise cloud assistant for the harder reasoning and drafting; and cheap cloud models for easy work that is not confidentiality-sensitive. You do not need all three on day one, but knowing the map keeps you from overpaying and over-exposing.
A note on the parade of specialized "AI for accounting" products. Many are a general model wrapped in an accounting-specific interface, which can be worth it for the workflow integration and the domain guardrails — but evaluate them on the same Module 2 criteria as anything else: what happens to the data, is training disabled, what is the retention, is there a business agreement. A friendly accounting label does not exempt a tool from the confidentiality review, and some wrappers route your data through the underlying provider on terms you need to read.
Cost intuition and management
Cost is measured in tokens — those roughly three-quarter-word chunks from Module 1 — and providers price per token of input and output, with frontier models costing meaningfully more per token than smaller ones. You do not need to track tokens obsessively; you need enough intuition to avoid waste, and a few habits deliver it.
Match the model to the task, the single biggest lever. Running every request through a frontier model is like sending a partner to do data entry — it works and it is wasteful. Route easy, structured work to cheap or local models and reserve the frontier model for tasks that actually need its reasoning. Across a firm's volume, this is the difference between trivial spend and a surprising bill.
Do not send more context than the task needs. Pasting a 200-page document to ask about one section costs tokens for all 200 pages and, per the Module 1 point about attention across long context, can worsen the answer. Send the relevant portion. Shorter, focused prompts are cheaper and often better.
Batch similar work. When you have fifty transactions to categorize, one well-structured prompt handling all fifty is cheaper and faster than fifty separate conversations, because you pay the instruction overhead once. Batching is a natural fit for the structured, cheap-model tasks and compounds the savings.
Be realistic about scale. For a typical firm, the direct token cost of well-managed AH-assisted work is small relative to the time it saves. The spending risk is not the per-request cost; it is the unmanaged pattern — everyone on frontier models, whole documents pasted for one-line questions, no standard tool. Manage the pattern and the cost takes care of itself. The far larger economic question is return, which we come to below.
A structured way to evaluate a new tool
New AI products arrive constantly, and staff will bring them to you asking "can we use this?" A firm needs a repeatable evaluation rather than an ad hoc reaction each time, and the criteria fall into three groups you can run in order.
Confidentiality and data handling — the gate. This group is a pass/fail gate, not a scoring factor, because a tool that fails it is disqualified for client work regardless of how good it is. Run the Module 2 checklist: is there an enterprise tier with training on your inputs disabled; what is the retention; is there a business agreement and, where relevant, the data-protection terms; who are the subprocessors and where is data processed. A tool that cannot clear this gate can only be used, if at all, on non-confidential material with de-identification — and most firm work is confidential, so the gate does most of the deciding.
Capability and fit — the value. Only for tools that clear the gate do you assess whether they actually help: does the tool do a job your firm has, does it do it well enough to save real time net of review, does it fit into how your staff already work, and is it meaningfully better than doing the same task through your existing approved general assistant. Many specialized products fail this test not on safety but on redundancy — they do something your general enterprise assistant already does, at additional cost and additional vendor risk. Be willing to conclude "this is fine, but we do not need it."
Cost and operational reality — the sustainability. What does it cost at your usage, who owns it internally, what happens to your data and workflows if the vendor changes terms or shuts down, and how hard is it to move off. AI vendors are numerous and not all will last; avoid building a critical workflow on a tool you could not replace. Prefer arrangements where your data and your prompt templates remain portable.
Running a tool through gate, value, and sustainability — in that order, stopping at the first failure — turns "can we use this?" from a debate into a short, consistent process. Name the person who owns it, per the policy below, so the answer is authoritative rather than whoever argued hardest.
A firm AI-use policy checklist
Module 2 argued that a written policy is how confidentiality scales from your careful habits to a whole staff. Here is a concrete checklist a firm can adopt and adapt. The least experienced person, on their busiest day, is who the policy actually protects — so it must be clear enough to follow without re-deriving the principles.
- Approved tools and tiers. Name the specific tools staff may use for firm work and the tier each is approved on. Using a non-approved tool for client work is a policy violation, not a judgment call.
- Data classification and routing. State plainly what categories of data may go into which tools. Client-identifying information, return data, and protected data go only into vetted enterprise (or local) tools per Module 2 — never a consumer tool. De-identification is expected even in approved tools.
- Mandatory human review. Require that a qualified professional review every AI-assisted output before it is used, delivered, or relied upon. The reviewer is responsible for the result. No AI output reaches a client or a filing unreviewed.
- No fabricated authority. Any citation to law, standards, or guidance produced with AI assistance is verified against the primary source before use. This is a specific, named rule because it is a specific, named risk.
- Cost-segregation and similar preliminary uses. Where AI produces a preliminary classification or triage (like the cost-seg intake), the policy states it is preliminary, is reviewed by a qualified professional, and does not replace the full professional engagement the work requires.
- Tool vetting ownership. Name who is responsible for evaluating and approving a new tool before anyone adopts it, using the Module 2 criteria. No one adopts a tool for client work unilaterally.
- Engagement-letter alignment. Ensure the firm's technology-use disclosure in engagement letters is consistent with actual practice.
- Incident handling. State what to do on a suspected data exposure — who to notify, how to contain it, how to document it. People make mistakes; the policy makes the response fast and correct.
- Training and acknowledgment. Require that staff who use AI for firm work understand its failure modes (hallucination, the confidentiality rules) and acknowledge the policy. This course is one way to meet that bar.
A one-page version of this is worth more than a ten-page version nobody reads. The goal is rules people actually follow.
Change management: adoption is the hard part
The technology is the easy part; getting a firm to use it well is the hard part, and it fails in predictable ways. Two failure modes dominate. The first is reckless adoption — staff pasting client data into whatever free tool they found, treating output as authoritative, skipping review because the answer "looked right." This is what the policy and the confidentiality training exist to prevent, and it is dangerous precisely because it produces fast, plausible, wrong work. The second is blanket refusal — a firm that bans AI outright and cedes the efficiency to competitors while staff quietly use consumer tools anyway, which is the worst of both worlds: no governance and no upside.
The path between them is deliberate, governed adoption. Start with a small number of low-risk, high-value workflows from Modules 4 and 5 — PBC lists, meeting notes, first-draft communications are good beginnings because the stakes are modest and the time savings are visible. Pick a champion who understands both the technology and the firm's work to shepherd it. Train staff on the failure modes before the features, because a person who understands hallucination and the confidentiality rules will use any tool more safely than a person who only learned the features. Make the review step a visible, expected part of the workflow rather than an afterthought, so it does not erode under deadline pressure. And normalize saying "this is an AI draft I reviewed," so the review is culturally expected rather than quietly skipped.
Resistance is often reasonable and worth hearing. Staff worry about accuracy — correctly, which is why review is mandatory. They worry about their roles — and the honest framing is that AI handles the tedious first-pass transformation so professionals spend more time on judgment, review, and client relationships, which is where their value was always concentrated. Address the concerns directly rather than mandating from the top; adoption that people understand sticks, and adoption that is imposed gets worked around.
Measuring return honestly
The reason to do any of this is return, and it should be measured rather than assumed in either direction. The primary return is time: hours saved on drafting, transformation, and triage that redeploy to higher-value work. Measure it concretely — time a workflow before and after AH assistance across a few real instances, and be honest that the AI time includes the review step, because a workflow that saves drafting time but adds equal review time has not saved anything. Most of the workflows here show a genuine net gain because reviewing a good draft is faster than creating one from scratch, but verify it for your own work rather than taking it on faith.
Watch the second-order effects too. Faster turnaround can improve client responsiveness and let you take on more work without adding headcount. Better-structured first drafts can raise the floor on quality, especially for less experienced staff who now start from a competent draft. And there is a real cost to get wrong: an unreviewed hallucination that reaches a client is a loss that dwarfs the time saved, which is why the review step is not overhead to be trimmed but the thing that makes the return real. Measure return net of review, and count the avoided-error value of review as part of the system working, not as friction.
A maturity path
Firms tend to move through recognizable stages, and knowing the path helps you place yourself and choose the next step rather than trying to leap to the end.
Stage one — individual experimentation. A few people use AI on their own, unevenly, often on consumer tools. Value is real but random, and confidentiality risk is highest here because nothing is governed. The right next step is not to stop; it is to govern — pick an approved tool and write the basic policy.
Stage two — governed basics. The firm has a vetted enterprise tool, a one-page policy, and a handful of standard workflows people actually use. Confidentiality is under control and the time savings are visible. Most firms should aim to reach and consolidate here before going further; it captures much of the value with manageable risk.
Stage three — embedded workflows. AI assistance is a standard, expected part of specific processes, with prompts standardized, review steps built in, and cheaper or local models routed to the easy work. The firm is measuring return and refining. This is where the technology stops being a novelty and becomes infrastructure.
Stage four — integrated and improving. AI is woven into the firm's tools and processes, staff are fluent, the policy is a living document that keeps pace with changing provider terms, and the firm evaluates new capabilities against a clear standard. Few firms need to rush here, and no firm should skip the earlier stages to get here — the governance and the habits built in stages one through three are what make the advanced stage safe rather than reckless.
The path is deliberate, not fast, and that is the point. A firm that reaches stage two with real confidentiality controls and a genuine review culture is in a far better position than one that reached stage four by pasting client data into consumer tools.
Recap, tied to the learning objectives
You began this course unable to reliably predict when an AI tool would help and when it would mislead you. You can now: an LLM generates likely language rather than retrieved truth, remembers only what is in its context window, varies because it samples, and fails through fluent, confident hallucination — most dangerously in citations, specific figures, and clean-but-wrong reasoning (Module 1). You can now state and apply the confidentiality guardrails that keep firm use consistent with your professional obligations: client data goes only into vetted enterprise or local tools with training disabled and retention controlled, minimized and de-identified even then, under engagement-letter disclosure and firm policy, with a qualified professional responsible for every output (Module 2). You can write a professional prompt using role, context, task, and format, strengthened with examples, grounding, reasoning, and guardrail phrasing (Module 3). You have ten concrete workflows with copy-paste prompts, each with its review step, and you know which deserve a frontier model and which run fine on a cheap or local one — including the cost-segregation triage, which produces a preliminary bucketing a professional verifies and which never substitutes for the full engineering-based study a filing requires, and the research workflow, whose iron rule is that authority is verified against the primary source before use (Modules 4 and 5). And you can now build this into a firm on a budget, with a policy checklist, honest cost and return management, and a maturity path that treats governed adoption as the goal (Module 6).
The through-line has been one idea in many forms: AI is a fast, fluent, unreliable assistant, and the profession lives in the review. Used inside that boundary, it is a genuine multiplier of a professional's time and reach. Used outside it, it is a confidentiality breach and a fabrication engine wearing the costume of competence. You now know the difference, and you know how to build your practice on the right side of it.