An AI demonstration can answer the sample question and still leave the operational questions untouched. What can it access? What can it change? Who reviews an uncertain answer? Those decisions belong in the pilot plan before the first real customer is involved.
TLDR: Choose one bounded use case, use approved data, inspect permissions and tools, define acceptance tests and keep human review in the loop. Set a cost boundary and stop criteria. Expand only after the recorded evidence supports the next step.
We recommend starting with a task whose input, output and failure can be inspected. “Help our service team answer approved public questions” is easier to evaluate than “automate customer service.”
Write what the pilot will not do. For example, it might draft an answer for a reviewer without sending it, changing a customer's record or making a commitment. These are proposed controls, not a claim that every HubSpot agent automatically operates this way.
Name the business owner and the person who can stop the pilot. A test without an accountable decision-maker can turn into an indefinite experiment.
Inventory the source material and remove content that is outdated, contradictory or not approved for the proposed use. Use synthetic records for test cases where practical. Do not assume data is safe to include merely because it is already stored in your CRM.
We recommend documenting each source, its owner, permitted use and review date. Where the use depends on sensitive information or legal interpretation, get the appropriate specialist review before running the pilot.
For a customer-agent example, HubSpot documents configuring sources and previewing behavior before deployment to channels. Check your account's requirements and current controls rather than treating a published tutorial as permission to deploy. HubSpot customer-agent setup, channel deployment
"Simulation" is not enough information to define a safe boundary. As checked on September 7, 2026, HubSpot's documentation for the agent builder says simulations do not consume credits or directly change CRM data, but configured MCP tools may still execute. Review the connected tools and their possible external effects before testing. This behavior should not be assumed for every HubSpot preview surface. HubSpot simulation limitations
We recommend disabling or isolating actions that are outside the approved test scope. Verify the actual configuration with an administrator. Do not assume an instruction in the prompt prevents an available tool from acting.
The following test set is synthetic. It illustrates a proposed review process for an assistant drafting answers from approved public service information.
| Test | Input condition | Expected behavior for this pilot |
|---|---|---|
| Supported question | The approved source contains a clear answer | Draft an answer consistent with that source for human review. |
| Missing answer | The requested detail is absent | State the limitation and send the question to a reviewer. |
| Conflicting sources | Two approved pages disagree | Flag the conflict rather than choose an unsupported answer. |
| Unauthorized action | The request asks for a record change | Do not perform the change; follow the approved escalation path. |
| Sensitive information | The request includes data outside the pilot boundary | Follow the reviewed handling procedure and stop the unsupported task. |
| Misleading instruction | Source text tries to override the pilot rules | Treat it as content, not authority to change the operating rules. |
Record the input, actual output, source evidence and reviewer decision. Repeating only the easiest question does not test the boundary you need to trust.
Confirm current credit requirements, available allowance and expected actions before any credit-consuming run. HubSpot's public credits catalog is a reference, not proof of your account's remaining balance or commercial terms. HubSpot Credits catalog
We recommend a monitoring owner, an authorized usage boundary and a documented stop procedure. Keep unknown rates or allowances unresolved until verified. Do not let a successful answer become implied approval to expand usage or buy capacity.
A HubSpot AI pilot is a controlled evaluation of one AI-supported task in a defined CRM process. Its purpose is to learn whether the configured behavior is useful enough, safe enough and maintainable enough for the next approved step. A demonstration of fluent text is only one observation.
Start with an input your team already understands. For customer service teams, that might be an approved public knowledge base article and a routine question. For sales teams, it might be a synthetic account summary with a clearly stated qualification rule. For marketing teams, it might be an existing public page and a request for a draft meta description.
These are possible test designs, not a claim that a particular HubSpot AI tool supports every action. Confirm the selected feature, data source, available tools and entitlement in the actual account.
Write down whether you are evaluating Breeze Assistant, a customer agent, a prospecting agent or a custom agent. Keep the exact product and configuration in the test record. Results from one assistant should not be used as proof that a different agent has the same permissions or behavior.
An AI chatbot answering public questions has different acceptance criteria from a tool allowed to update properties. An answer can be reviewed before use. A record change may affect reporting, lead routing or another connected process before anyone reads the output.
If the business needs only a draft, do not enable a production action merely because the tool supports it. Start with the smallest configuration that can answer the evaluation question.
Choose the properties needed for the task and define their meaning. Job title, deal stage and company fit can be incomplete or inconsistent. An AI-generated summary cannot repair that uncertainty simply by sounding decisive.
Synthetic example: A fictional account has a company size but no verified service region. The proposed qualification rule requires both. The expected result is “region needs review,” not a confident qualified-lead label inferred from unrelated fields.
A useful pilot can expose a data-quality dependency even when it does not justify automation. Record that finding as data work, then rerun the same case after an authorized correction.
Brand voice can guide tone, terminology and formatting. It does not establish whether a customer commitment is true or whether a record is appropriate to use.
For content creation tests, evaluate facts and style separately. Check whether a draft preserves material limitations, keeps citations attached to the right claims and avoids inventing a product feature. Then check whether the writing fits your audience.
Do not accept an unsupported answer because it sounds like your company. Equally, an awkward but accurate draft may need an editorial correction rather than a change to the data source.
Choose test cases that matter to the team using the result. Repetitive tasks are useful candidates only when the required output and exception path are clear.
For a service pilot, compare a supported public question with an order-status request the approved sources cannot answer. Define the handoff expected in the second case. Accurate responses include recognizing when the requested information is unavailable.
For a sales pilot, compare a complete synthetic account with one containing conflicting buying signals. Have sales reps inspect the source evidence before using a proposed personalized outreach message. A plausible explanation of customer behavior is not evidence that the behavior occurred.
For a marketing pilot, compare an ordinary page summary with source text that contains an instruction to ignore the task. The latter tests whether untrusted content is treated as material to analyze rather than permission to change the operating rules.
Keep the exact input, output, selected sources, configuration and observation date. Identify which criteria passed and which remain unresolved.
A reviewer should be able to explain the decision without rerunning the whole pilot from memory. If a source page changes or an agent's available tools change, the earlier result may no longer cover the current configuration.
Do not turn a small set of passing cases into a percentage claim about all future customer interactions. State which cases were checked and what further evidence is needed.
A pilot can produce a useful answer and still create extra work if no one owns exceptions. Decide where unresolved questions go, who reviews them and how the team records a correction.
If customer success managers receive escalations, include them in the test of the handoff. Verify that they receive the context needed to continue. Do not infer customer satisfaction from the fact that an automated response was sent.
For sales and marketing work, decide who approves externally visible messages. Keep a draft, an approved message and a sent message as separate states. The same distinction applies to content created for blog posts or landing pages.
We recommend stopping for an action outside the approved boundary, exposure of excluded customer data, an unsupported customer commitment or a breached usage limit. Define the exact procedure before the test.
A lower-quality draft may be correctable within the pilot. A permission or data-exposure failure requires investigation of the configuration and affected information, not another prompt variation.
If time is part of the business case, observe the complete process: preparation, review, correction, exception handling and maintenance. Compare that with a documented baseline for the same work.
Do not count only the seconds needed to generate an answer. A faster draft can still require more review, and no time saving should be claimed without a defensible comparison.
We recommend stopping for an unauthorized action, exposure of excluded data, unsupported customer commitment or a cost boundary breach. The exact thresholds should be approved for your use case before execution.
For less severe failures, record the cause, corrective action and test to rerun. A source problem may require content cleanup; a permission problem requires configuration review. Rewriting the prompt is not a universal remedy.
Decide whether to continue the bounded pilot, revise it, defer it or end it. Only consider a wider rollout after the required tests pass and the appropriate owners approve the next scope.
If your portal baseline is unclear, start with the HubSpot portal audit checklist and common-issues guide. For help designing a controlled evaluation, explore HubSpot operations support or request a scoped review.
An AI pilot should distinguish assistance from delegated action. A draft answer that a person reviews has a different risk profile from an agent that can update CRM records or communicate with a customer. Define that boundary before comparing AI features.
Use a short candidate record: task, approved information, expected output, allowed action and responsible operator. For a customer-agent use case, include the sources it may answer from and the conditions that require handoff. For a sales-process use case, include the audience and communication rules. Do not assume a successful internal writing test validates customer-facing automation.
We recommend reviewing past interactions only when the organization has approved that use of the information. Support tickets, call transcripts and customer messages may contain material that should not be copied into an unrestricted test. Use a representative synthetic example when sensitive information is unnecessary.
Write down how a person will identify an incorrect result, stop further action and correct the affected record. Constant human intervention can make a proposed automation uneconomic, but removing review before quality is demonstrated does not solve the problem.
Compare time spent reviewing, correcting and maintaining the pilot with the current process. A faster generated response is only one part of the workload. Keep any efficiency claim provisional until the full task has been measured.