AI in Accounting: Lessons from the HSP ChatGPT Case

AgentSunrise
AI in Accounting
ChatGPT for Accounting
AI Implementation
Enterprise ChatGPT

In brief: On August 7, 2026, OpenAI published the HSP GRUPPE case study: the shared ChatGPT Enterprise workspace across professional firms reached 84% weekly activity. The strength of the case is not any single feature, but the combination of rules, training, shared AI agents, and mandatory expert review. For Russian businesses, this is not a promise of ROI, but a verifiable implementation model.

This material is intended for leaders of accounting, finance, tax, audit, and legal functions. We examine the organizational mechanics of the case and pilot design; we do not provide tax or legal advice and do not claim that German processes, metrics, or data-handling rules automatically apply in Russia.

Updated: August 10, 2026. Facts and product boundaries are set as of this date.

Contents

What happened on August 7, 2026

OpenAI published the HSP GRUPPE case study — a network of legally independent tax, audit, and law firms. The shared ChatGPT Enterprise workspace of HSP and Kanzleipakt covered 81 organizational groups. This is an important detail: the metrics refer to the shared workspace, not just one HSP firm.

From February 1 through July 14, 2026, OpenAI recorded 755 weekly active users, 913 unique users and more than 500,000 messages. HSP used the system for tax and legal research, client communication, financial analysis, and knowledge sharing.

The organization did more than provide access to chat. It held monthly AI forums, shared best practices, and turned repeatable workflows into shared Agents. For example, AI Client Communication helped prepare the structure and first draft of messages, while Booking Assistant SKR03 & SKR04 helped sort out individual classification questions. Professional review and final responsibility remained with the subject-matter specialist.

The HSP GRUPPE website confirms the organizational context: it is a cooperative network of legally independent tax, audit, and law firms. That is why the case is interesting not only for a large centralized corporation. It shows how to standardize knowledge and practices across autonomous units without eliminating local accountability.

What results HSP reports

Below are only the figures published in OpenAI’s customer story. They are Measured for what the source reported and measured, but not an independent market assessment.

Metric Value and time window Type of evidence
Weekly activity 84%; 755 weekly active users shared workspace telemetry
Unique users 913 over six months shared workspace telemetry
Messages more than 500,000 from February 1 to July 14, 2026 shared workspace telemetry
Higher productivity 98.6% of respondents employee self-assessment
Improved work quality 84.6% of respondents employee self-assessment
Weekly time savings 95.9% of respondents; 63.5% saved at least two hours, 25.7% at least five employee self-assessment
Real estate investment analysis from about nine hours to two hours for one partner named example, not a controlled experiment
Additional annual capacity about 40,000 hours conservative internal HSP scenario
Revenue potential about €3.8 million per year theoretical scenario, not realized or guaranteed revenue

Here it is important not to mix four levels:

  1. Usage telemetry shows that people actually used the system.
  2. Survey shows employee perceptions, but it depends on the sample and the wording of the questions.
  3. Named example shows that a specific workflow is possible, but it does not show the average effect across the network.
  4. Capacity model turns assumptions into an internal scenario; this is still not a financial result.

What the case proves and what it does not prove

The case proves that a specific network of professional firms was able to achieve high use of a shared workspace and build repeatable workflows. It also documents management elements: rules, training, shared Agents, small-group piloting, and preservation of professional responsibility.

But the publication does not disclose the denominator or the survey questions, the distribution of activity across 81 groups, a control group, result quality before and after implementation, or the full 40,000-hour calculation. So it cannot be used to make the causal claim that “ChatGPT increased productivity by 98.6%.”

Can conclude Cannot be concluded
755 users were active weekly in the specified workspace 84% activity means 84% growth in productivity
most respondents reported saving time each employee will save at least two hours
HSP standardized separate practices into shared Agents any agent produces a correct tax or legal answer
one partner reported reducing analysis time from nine hours to two the average time for all similar tasks will decrease by 78%
HSP calculated a theoretical potential of €3.8 million the company has already generated that revenue
the organizational model deserves review the German result is being applied to Russian data, systems, and rules

Official OpenAI documentation on workspace analytics draws the same line: analytics describes adoption and activity, but does not change access rights and is not an audit log. Exported reports should be treated as identifiable organizational data and subject to the organization’s own access, retention, and data storage rules.

Why a license is not enough

Providing access creates a technical ability, but not an organizational capability. In professional services, value appears when an employee knows not only “what to ask,” but also:

  • which data may be shared;
  • which source is considered acceptable;
  • where a second review is required;
  • who signs off on the result;
  • which correction is considered material;
  • when the scenario must be stopped.

HSP was not starting from scratch: the network had invested for more than two decades in digitalization, process standardization, and quality management. That explains why successful individual practices could be turned into shared workflows. AI accelerates a mature process and just as quickly scales its uncertainty if there are no rules.

That leads to a counterintuitive takeaway: the first implementation artifact should not be a prompt, but a process map. It shows the inputs, owner, sources of facts, permitted action, review point, and result log.

Which tasks are suitable for the first rollout

It is safer to start with reversible, verifiable preparation steps. The closer the action is to posting an entry, issuing an official opinion, creating a payment instruction, or sending a message to a client, the stricter the controls and authority should be.

Level Example AI role Required control
Low structure an internal note, create a list of missing documents draft and classification spot check by the process owner
Medium prepare a client letter based on an approved template draft with sources line-by-line review by a specialist before sending
Medium compare financial scenarios using specified formulas calculation and explanation reconciliation of inputs, formulas, and result
High classify an unusual accounting question hypothesis and options the decision is made by a qualified employee
High interpret a tax or legal rule research and structure arguments review of the current primary source and specialist sign-off
Not allowed for autonomous launch file reports, change an entry, issue a final opinion do not automate without a separate approved workflow technical restrictions, separation of duties, audit

This is not a universal legal classification. Each company should separately define data classes, confidentiality, local requirements, access rights, and the cost of error together with the relevant responsible experts.

The PROFI framework: how to manage AI in professional work

AI Dawn offers the framework PROFI: Process → Risks → Responsible → Facts → Measurement. This is an original management model for a pilot, not a ChatGPT feature or an HSP research finding.

1. Process

Choose one repeatable stage, not the entire job title. A good candidate has a clear input, an observable result, and a history for comparison: drafting a response, checking whether a package is complete, or initially grouping requests.

Establish a baseline: how long the stage takes, how many results are accepted without major rework, what errors occur, and where the client waits. Without a baseline, “it got faster” will remain just a feeling.

2. Risks

Separate data and actions by class. For each class, define permitted tools, storage, access rights, retention period, and prohibited operations. Separately define what the system must not do even when it gives a confident answer.

Do not rely on the generic slogan “the data is protected.” HSP explicitly notes that the use of client information is governed by its internal requirements for data protection, confidentiality, and governance. A product’s technical feature does not replace the organization’s policy.

3. Responsible

Assign the person who reviews and accepts the result. The phrase “human in the loop” is too vague: you need the role name, response time, escalation criteria, and the authority to stop the scenario.

For a client letter, the owner may be a consultant; for classifying an entry, an accountant; for changing a workflow, the process owner and information security. AI does not acquire professional accountability through a good prompt.

4. Facts

Require sources, date, version, and traceability. An answer without a citation may be useful as a hypothesis, but not as the basis for a professional decision. For a corporate knowledge base, record which document was used and when it was updated.

Official OpenAI documentation on governance separates interactive analytics, aggregated reporting, and the Compliance API for auditable records. This is a useful architectural boundary: a dashboard for adoption should not be presented as an evidence log.

5. Measurement

Measure three layers separately:

  • adoption: who uses the system and how often;
  • quality: what was accepted, corrected, or rejected;
  • business outcome: what changed in turnaround time, throughput, error cost, or customer outcome.

Mixing layers leads to a false conclusion: “lots of messages means there’s impact.” The right question is different: what share of the work did the specialist accept, how many material errors were corrected, and what changed versus the baseline.

30-day pilot in shadow mode

Shadow mode — a mode in which the AI prepares output in parallel with the current process but does not send it to the client or change the accounting system. Thirty days is an estimated Estimated organizational horizon, not a universal deadline.

Days 1–5: freeze the task

  1. Choose one process and 30–100 historical or current cases.
  2. Define the baseline, data classes, and prohibited actions.
  3. Assign a review owner.
  4. Set the acceptance criterion before testing begins.

Days 6–15: build the parallel stream

The system prepares drafts, but the employee works through the normal process. After completion, the owner compares the options using one rubric: completeness, correctness, sources, material corrections, review time, and potential risk.

Days 16–25: verify repeatability

Repeat the test on new cases and separately review edge cases: incomplete data, conflicting documents, a new regulation, an unusual client request. Do not retroactively improve the criteria just to save a pretty demo.

Days 26–30: make a decision

There are four honest outcomes:

  1. expand to similar tasks;
  2. keep it as a drafting tool;
  3. narrow the scenario or strengthen the sources;
  4. stop the rollout if oversight costs more than the savings.
process → constraints → AI draft → specialist review → facts → metrics → scale decision

Which metrics separate activity from impact

Layer Metric How to measure What it does not prove
Adoption weekly active users users with activity / authorized users quality and savings
Adoption scenarios per user completed scenarios / active users the usefulness of each scenario
Quality acceptance rate accepted results / reviewed results absence of hidden errors
Quality substantial correction rate results with material edits / reviewed results causal financial impact
Quality source completeness answers with acceptable sources / answers where a source is required source freshness without verification
Risk policy incident rate policy violations / processed cases full detection of violations
Outcome median cycle time median time from intake to accepted result revenue without connection to workload
Outcome realized capacity hours actually redirected to validated work automatically generated profit

Do not start with a financial model. First confirm quality and safety at the task level, then process change, and only after that connect freed capacity to actual utilization or revenue.

Frequently Asked Questions

What did the HSP GRUPPE case show?

The case showed high usage of a shared ChatGPT Enterprise workspace and positive employee self-assessments across a network of professional firms. These are vendor and participant data, not independent proof of ROI or of the result’s transferability to Russia.

Can ChatGPT be assigned the final tax or accounting decision?

No. In the HSP case, professional review and final responsibility remain with the tax, legal, or accounting specialist. High-risk actions require separate authority, technical controls, and auditability.

Where should AI implementation in accounting start?

Choose one limited process, document the baseline, data classes, owner, factual sources, and acceptance criterion. Then run a shadow-mode pilot without automatic client delivery or changes to the accounting system.

What metrics are needed besides user activity?

Measure the share of accepted results, the number of material corrections, cycle time, source completeness, policy violations, and actually freed capacity. Weekly active users show adoption, but they do not prove productivity.

Can HSP’s numbers be applied to a Russian company?

No. Processes, data, systems, regulation, culture, and baseline all differ. HSP’s numbers are useful as a hypothesis for your own measurement, but not as a promise of results.

How AI Dawn helps build a controlled operating model

AI Dawn connects AI not with abstract “accounting digitalization,” but with a specific process and acceptance criterion. For a professional workflow, the team can take four steps:

  1. Describe the process, baseline, data classes, risks, and the role of the final reviewer.
  2. Prepare the data and corporate knowledge base with sources, versions, and access rights.
  3. Build an MVP agent or workspace, add an action log, and run a shadow-mode test.
  4. Integrate a validated workflow with 1C, ERP, a browser-based interface, or an internal system, and train the team.

A safe first step is to choose one process, its current baseline, data sources, constraints, and acceptance criteria. This lets you test value before a large-scale integration and avoids promising timelines, savings, or financial results upfront.

Discuss the project

Conclusion

The latest HSP case study matters not because of the 500,000 messages or the theoretical €3.8 million. It shows a sequence: a standardized process, a shared workspace, data rules, knowledge exchange, shared Agents, and professional accountability.

For Russian businesses, the useful thing to copy is not a specific tool or someone else’s percentage, but the management framework. The PROFI framework forces you to define the process, risks, owner, facts, and measurement before scaling.

The first practical step is to take one repeatable stage and 30–100 cases, run a 30-day shadow-mode test, and agree on a stop/go criterion in advance. If the system increases the share of accepted outputs without increasing material errors and truly shortens the cycle relative to the baseline, the scope can be expanded. If not, the negative result will protect the company from an expensive rollout based on a flashy case study.

Request an audit

Share your contact details and we will follow up.

← All articles

Comments (0)

Loading comments…

Leave a comment
No registration required

Book a strategy call
for agentic operations

Tell us which workflow you want to improve. We will map feasibility, risks, and the fastest MVP path.

By submitting, you agree to our privacy policy

Contacts

Global Operations

Serving U.S. clients remotely
with private cloud and on-prem options

Strategy calls by request

We respond after reviewing your workflow context.

lamooof@gmail.com

For partnership inquiries

Have a proposal?

Write to us in messengers

© 2025 AgentSunrise