Protego field desk
AI Security9 min read

Chatbot Attack Probes: How We Protect Our Site Guide

We investigated three injection-style probes in our chatbot logs. Here is what the evidence shows, how the Site Guide works, and what we are improving.

I
Microsoft Cloud Solution Architect
A chat bubble passes through a bounded shield to public resource cards, illustrating the Protego Site Guide
A chat bubble passes through a bounded shield to public resource cards, illustrating the Protego Site Guide
chatbot securityprompt injectionsecurity investigationapplication securityProtego

Someone sent our chatbot {{7*7}}. It answered: “That's 49.”

That exchange looked alarming in the admin dashboard. Curly-brace expressions are familiar to anyone who tests for server-side template injection. Did the response mean someone had made our server execute their input?

We investigated on October 4, 2026. We reviewed all 14 stored exchanges, compared them with available request records, and traced both the historical and current chatbot code. We found three injection-style probes in one short burst. We found no evidence of secret disclosure, an authentication bypass, or server-side code execution in those exchanges.

There are two qualifications. The probes were recorded in July, not during a new October attack campaign. Their request metadata was consistent with local testing, so we cannot responsibly attribute them to an outside attacker. Our historical logging also has gaps.

This is a chatbot security case study about interpreting evidence and limiting what a chat interface can do. It is not a claim that Protego defeated a sophisticated intrusion.

The three probes we found

The suspicious inputs appeared on July 5, 2026 UTC, within a burst lasting less than a minute. We have omitted addresses, session identifiers, ordinary visitor questions, and private operational details.

Probe typeWhat it was testingRecorded responseWhat we can conclude
SQL-style authentication bypassWhether input could change a database query or login decisionThe chatbot identified the pattern and offered security-learning guidanceNo database result or authentication change was shown
Template expression: {{7*7}}Whether a template engine would evaluate an expressionThe older chatbot answered “That's 49” and returned to cybersecurityArithmetic in a model response is not proof of template execution
Alternative template expression: <%= 7*7 %>Another template-evaluation syntaxNo assistant answer; the request was recorded as rate-limitedThis request reached the configured daily limit

These are recognizable injection probes. They are not enough to establish who sent them, whether the sender had malicious intent, or whether a vulnerability existed. A developer testing a local application can submit the same strings as an external scanner.

Our records contain metadata consistent with that local-testing scenario. We could not establish attribution, and we are not publishing an attacker identity or campaign claim.

Why “49” did not prove server-side template injection

Server-side template injection happens when untrusted input is interpreted as template code. An arithmetic expression is a useful initial probe because an evaluated result can reveal that interpretation. OWASP's template-injection testing guide documents this approach.

A chatbot adds another explanation: a language model can recognize an arithmetic expression and produce the answer as ordinary text. The same visible number can come from two very different execution paths.

We therefore examined the application code available from before the recorded requests. That version sent the visitor's message and recent conversation history to a language model. It streamed the model's text back to the visitor. It did not expose a shell, a template evaluator, or callable administrative tools to the model.

The reviewed path also did not insert the message into a database query. It stored the message as a log field and maintained a usage counter. Writing a string into a record is different from executing that string as a query.

The observed response is consistent with the model doing arithmetic conversationally. We could not recover the exact historical deployment identity, so source history and saved responses remain the evidence available to us. We did not find evidence of template execution.

There was a small behavioral deviation: the chatbot answered an off-topic arithmetic question before redirecting to security. That matters for product behavior, but it is not equivalent to gaining server access.

Injection probes and prompt injection are different questions

It is tempting to call every strange chatbot message a prompt-injection attack. That blurs the diagnosis.

A SQL-injection probe asks whether data becomes part of a database command. A template-injection probe asks whether data becomes template code. A prompt-injection attempt asks whether untrusted content can redirect a model away from its intended instructions or authority boundaries.

Our observed three-probe burst contained SQL-style and template-style inputs. We did not find a recorded “ignore your instructions and reveal secrets” attempt in the reviewed exchanges. We did use harmless instruction-override examples in a separate offline check of today's responder; those synthetic tests are not visitor attack statistics.

For the broader model-specific problem, see our guide to prompt injection in enterprise copilots. OWASP likewise treats prompt injection as a risk that requires controls beyond a well-written system prompt.

How the current Protego Site Guide is protected

The implementation changed in August. The public Site Guide now uses deterministic matching rather than a language-model API. Its job is narrow: help someone find a published guide, a learning resource, or the right security tool.

A message follows this simplified path:

The current Site Guide response path

Simplified successful path. Rejected or failed requests stop earlier.

  1. Step 1: Visitor message

    Input arrives as text, not an authorized command.

  2. Step 2: Bot verification

    The endpoint checks BotID before processing the message.

  3. Step 3: Usage limit

    Daily admission checks bound chat usage.

  4. Step 4: Rules and public resources

    Fixed replies or keyword matches select public content.

  5. Step 5: Text and links

    The response points to resources without invoking tools.

This diagram shows the intended successful response path. Rejected and failed requests terminate earlier. Logging supports investigation but is not a complete record of every request outcome.

The message cannot select a privileged action

The current responder matches words to fixed replies or public article metadata. A message requesting a password does not grant access to a password store. Asking it to run a command does not create a command-execution path.

It has no model-driven tools, shell, payment action, account-management action, or arbitrary visitor-directed URL fetch. The chat handler still performs ordinary application operations, including reading public content and writing usage records. The important boundary is that the message cannot instruct it to perform a different privileged operation.

If we add agent capabilities later, those capabilities will need their own authentication, authorization, restricted inputs, and appropriate human approval. OWASP's prompt-injection prevention guidance emphasizes limiting privileges and keeping authorization outside the model's judgment.

Conversation history is not an authority source

The current responder does not use client-supplied conversation history to decide its answer. A visitor cannot manufacture an earlier “administrator approval” message and have that become a server permission.

That design also limits conversational flexibility. The Site Guide can make poor keyword matches or miss a question. We accept that a narrow resource finder is less conversational than a general assistant; we still need to improve its relevance. Deterministic behavior does not automatically mean useful behavior.

Bot checks and usage limits reduce abuse

The route checks Vercel BotID and rejects requests classified as bots before processing the chat message. It also applies a daily usage limit. The historical final probe was recorded as rate-limited.

These controls have separate jobs. Bot verification addresses automated traffic; a usage limit bounds consumption. Neither proves that a human's message is benign, and neither establishes that an application is immune to injection. Vercel's BotID documentation describes the server-side verification mechanism.

Messages are treated as text

The reviewed widget and admin log view render message strings through React rather than inserting them as raw HTML. That helps preserve the distinction between displayed content and executable markup.

Links remain a separate security concern. Any application that turns text into clickable links must consider destination and scheme validation. Text rendering is one control, not a reason to ignore everything else in the response pipeline.

What our checks established, and what they did not

We manually reviewed all 14 stored exchanges. The only three injection-style probes were the historical burst described above. Recent request records available from the hosting platform matched three ordinary saved exchanges. A successful HTTP response indicates that a request completed; it does not establish whether an attack succeeded.

We also replayed eight harmless probes offline against the exact current response function. They included SQL-style text, arithmetic templates, instruction overrides, requests for secrets, and markup-like input. With an empty article fixture, all eight returned the expected fixed fallback.

That is a focused function check. It does not test every possible input, prove the entire website secure, or replace an end-to-end assessment of authentication, dependencies, rendering, and hosting controls.

The most important limitation is historical coverage. Usage counters contain six additional admitted requests without saved exchanges. A counter increments before a complete response exists, so the mismatch can include interrupted requests, processing failures, or missing log writes. We cannot reconstruct their contents or outcomes from a number.

Our finding is therefore specific: no successful compromise was demonstrated in the reviewed chatbot exchanges. We cannot turn incomplete records into an absolute historical all-clear.

What we are improving next

The investigation produced work beyond reviewing suspicious text. Our priorities are more reliable request-outcome logging, clearer separation of historical and current behavior, and further input, usage-control, and rendered-link hardening.

These are follow-up priorities, not claims that every improvement has shipped. Investigation records should make it possible to distinguish rejected, failed, interrupted, and completed requests without collecting unnecessary visitor information.

For another team reviewing its chatbot, the useful questions are concrete: Which version handled the message? Where did the input go? What operation could it trigger? What evidence shows that operation happened? What records are missing?

Our AI red-teaming guide provides a broader testing framework. Teams adding tool access should also review the boundaries described in our MCP server security guide.

Frequently asked questions

Was Protego's chatbot hacked?

We found no demonstrated compromise in the available exchanges. Three historical injection-style probes were present, with no recorded secret disclosure, authentication bypass, or server-side execution. Incomplete historical logs prevent an absolute guarantee.

Does answering a template expression prove code execution?

No. A language model can answer arithmetic as text. The investigator needs to trace whether untrusted input reached an actual evaluator and whether the claimed operation occurred.

Were the probes sent by outside attackers?

We cannot establish that. Their metadata was consistent with local testing. The payload pattern alone does not identify the sender or intent.

Can the current Site Guide run tools or change my account?

The current chat responder has no tool-execution or account-management capability. It matches a question to fixed guidance and public resource links. Other website tools are separate features with their own controls.

Does removing a language model make a chatbot completely secure?

No. It removes model-instruction following from this response path. Input handling, availability, authentication, dependencies, logging, and link rendering still require security work.

Evidence and responsible reporting

This case study is based on Protego's October 4, 2026 review of stored exchanges, current and historical application code, available hosting request records, and an offline response-function check. Public examples are limited to simple probe strings. We have not published visitor addresses, session identifiers, or unrelated conversations.

If you discover a reproducible issue, use the reporting information in our security.txt. Include the affected feature, the observed behavior, and a safe reproduction. An unexpected answer is a useful lead; the next step is establishing which boundary, if any, it crossed.

RD / 01

Reader desk

Discuss this guide

Rate the guide or ask a practical follow-up. Clear questions publish immediately; uncertain submissions wait for review.

Free download

AI Security Risk Assessment Template

Evaluate LLM and AI system risks with this structured assessment template.

No spam. Unsubscribe anytime.

Continue Learning

AI Security Engineer Roadmap

The fastest-growing specialty in security.

Start the Intermediate Path10h · 4 topics · 10 quiz questions
I

Microsoft Cloud Solution Architect

Cloud Solution Architect with deep expertise in Microsoft Azure and a strong background in systems and IT infrastructure. Passionate about cloud technologies, security best practices, and helping organizations modernize their infrastructure.

Share this article

Related Articles

Need Help with Your Security?

Our team of security experts can help you implement the strategies discussed in this article.

Contact Us