Skip to content
AgentSteamer ARTICLES / WIKI
AllArticles
Knowledge Base
Scene Main site

How Your Data Leaves the Company: The Six Lines of Defense Every Enterprise Agent Needs

模釜 · AgentSteamer 文章中关于企业 AI Agent 数据安全六道防线的配图

When enterprises choose an AI Agent, the real question isn’t “will it secretly transmit data,” but “does it have the ability to transmit data, and can you see what it transmitted.” The former depends on the vendor’s integrity and can’t be written into acceptance criteria; the latter is a verifiable, auditable technical fact that can be written into a contract.

Last month, during a vendor selection defense at a manufacturing client, the CISO across the table asked: “Will your platform secretly transmit our data out?” I said the question was backwards. Not out of guilt, but because with the “will it” framing, the answer always depends on the integrity of the company across the table, and integrity can’t be written into acceptance criteria. Three months ago, if someone had asked Zhipu “Will ZCode upload my project,” they probably would have said no too. The only two things you can ask are: does it have the ability to transmit data out, and can you see what it transmitted. This article covers those two things: first a review of several incidents that actually happened this year, then six requirements enterprises should put to Agent vendors.

One-Minute Overview

  • On September 18, 2026, reverse engineering analysis of Zhipu’s ZCode revealed that as long as a user is logged into an account, it packages the entire project along with historical modification records, encrypts them, and uploads them. The two seemingly relevant toggles in the local interface can’t stop it even when turned off; the decryption key exists only on the vendor’s side, and the package is generated on your own computer but you can’t open it yourself. That package was 313MB, containing roughly 42,000 files, of which 86.6% were the project’s historical modification records.
  • Similar incidents have occurred at least three times this year. xAI’s Grok Build uploaded entire projects to Google Cloud Storage, including files users had explicitly said “do not read” and unredacted passwords. Claude Code was directly reported by the Ministry of Industry and Information Technology’s NVDB for security backdoor risks.
  • These three incidents share one thing in common: the source of risk isn’t an attacker, it’s the company providing the tool itself.
  • All current Agent security frameworks defend against external attackers. They’re built on a default premise: the vendor and the user are on the same side. The exfiltration channels above happen to sit right in the blind spot of that premise—check them one by one against any framework, and not a single alarm triggers.
  • The real question enterprises should ask isn’t “is this Agent safe,” but “do the things I’ve authorized and the risks I bear balance out.”
  • AgentSteamer is an enterprise-grade AI agent platform developed by Shanghai Immersivalley Information Technology Co., Ltd. It supports full private deployment and isn’t tied to any specific large model.

I. First, a Review of Three Exfiltration Incidents, and One Type of Service

ZCode: A 313MB Package, and Two Toggles That Won’t Turn Off

On September 18, tech blogger ferstar noticed something off with disk usage while inspecting a local directory. He found a 313MB encrypted file he couldn’t open, but the accompanying file manifest showed roughly 42,000 files inside, over 80% of which were the project’s historical modification records. Not the project itself, but the entire history of this project from birth to the present.

Looking further, the problem wasn’t just the size. In this packaging process, the history records directory was cleared for inclusion before key filtering and size limits were applied. That means the filtering for password files like .pem and .key, and the 1MB per-file size cap, had no effect whatsoever on the contents of the history records directory. Any passwords and keys ever committed and later deleted were carried off as-is.

The bigger problem was the key. The client first requests upload credentials and a public key from the vendor’s server, completes compression and encryption locally, then bypasses the vendor’s own business server to upload directly to Alibaba Cloud’s object storage, which then calls back to the backend. The public key is the lock, the private key is the key; the lock is issued temporarily by the server, and the key is kept only on the server side. In other words, this package is generated on your own computer, but you can’t open it—only the vendor can see what’s inside.

There are two seemingly relevant options in the interface, one called “Optimize Experience” and one called “Repository Snapshot Index.” After checking the code logic item by item, it was confirmed: the former only controls whether data is used for model training, the latter only controls whether the server builds a retrieval index after receiving the data. Turn both off, and local packaging and uploading continue as usual. The component responsible for snapshots and uploading is loaded unconditionally when the software starts, with the only prerequisite being that you’re logged in. ferstar tried manually deleting the file pending transmission. Half an hour later, ZCode regenerated one. The file’s failure retry count at that point was 564.

Zhipu apologized that same night, attributing the cause to the “code repository indexing” feature, explaining that Repo Wiki might trigger uploads when generating pages, that data is immediately destroyed after cloud generation and not retained, that the feature was enabled by default in its early rollout, and that it has now been fixed. It also promised to open-source the code repository soon, bring in third-party evaluation, and gave all users a quota reset as compensation. The response was quick, but what it explained wasn’t the same thing the community was asking about. Indexing locally, snapshotting locally, rolling back locally—all technically feasible, and some tools on the market do exactly that. That leaves an unanswered question: if it can be done locally, why send the entire project to the cloud?

One more detail surfaced. ZCode v3.12.2’s changelog, dated September 16, two days before ferstar’s post, included an entry reading “optimize memory usage of repository snapshot upload.” It’s hard for an engineering team to optimize memory for an accidental behavior. That changelog entry was deleted after the incident gained traction.

Grok Build: Uploading Even the Files You Said “Do Not Read”

In July of this year, independent security researcher cereblab conducted a full packet capture analysis of xAI’s coding Agent tool Grok Build, publishing all evidence and reproduction steps. The results were more direct than ZCode’s: entire projects were packaged and uploaded to Google’s cloud storage service, with the upload scope covering all files, including those the user had explicitly said “do not read” in conversation. In a 12GB test project, confirmed file volume exceeded 5GB when the packet capture was interrupted. Password and key files in the project were likewise uploaded as-is, without any redaction. When users turned off the “Improve Model” option in settings, uploading continued as usual. What was turned off was only training authorization, not whether code left the machine.

Claude Code: The MIIT Notice

This one isn’t a community leak, it’s an official notice. The Ministry of Industry and Information Technology’s National Vulnerability Database (NVDB) monitored and found that the AI coding tool Claude Code has security backdoor risks of serious severity. Its built-in monitoring mechanism transmits sensitive information such as user location and identity identifiers back to remote servers without user consent. Affected versions are 2.1.91 through 2.1.196. NVDB’s recommended actions were specific: immediately uninstall or upgrade development terminals with the above versions installed, and strengthen outbound access control and traffic monitoring for development tools within core business network segments to prevent unauthorized exfiltration of sensitive data.

Before this, the community had made an earlier discovery: Claude Code polls the server once per hour for remote configuration, and the configuration items include multiple control switches that can force-quit the program and bypass user permission prompts, all taking effect in the background without requiring the user to actively update. It also reads environmental signals such as the user’s proxy, gateway address, and China timezone. An Anthropic engineer later confirmed it was a proactive experiment targeting account abuse prevention and anti-distillation.

AI Relay Stations: A Warning from the Ministry of State Security

The fourth item isn’t a product, it’s a type of service. “AI relay stations” are proxy layers between users and model vendors’ official services, integrating various models’ APIs into a single platform and providing it to users. They’re convenient and cheap—one entry point can call several vendors’ models, and can even bypass network access, official authorization, and cross-border transmission restrictions.

The Ministry of State Security listed four categories of risk:

  1. Data running naked: Data submitted by users is retained on the relay station’s servers; some platforms lack proper encryption and control mechanisms, and some even intercept it privately and resell it to other model vendors for training.
  2. Model downgrading: Passing off low-spec models as high-end ones, cutting compute, disabling validation, producing highly deviant outputs.
  3. Malicious implantation: Some relay stations hide backdoors, steal account keys and cloud credentials, and even implant remote control programs.
  4. Data leaving the country: Without obtaining data export compliance qualifications or completing security assessment procedures, they transmit user input to overseas servers without authorization.

Behind these four categories of risk is the same structure: you’ve inserted a third party in the middle that you can neither see nor control.


II. Why Existing Security Frameworks Can’t See This

Over the past year, Agents have been granted more permissions than any previous type of software installed on personal computers. It can read all files in a project directory, autonomously execute command-line operations, while maintaining a constant connection to the vendor’s server, and can also receive remote configuration updates in the background. Before this, almost no consumer-grade software satisfied all four of these conditions simultaneously.

Rules are indeed being filled in quickly. At the end of 2025, OWASP released the first top-ten risk list for autonomous AI Agents; in January 2026, Singapore introduced the first governance framework for autonomous AI Agents, requiring each Agent to carry a verifiable digital identity; in February, the U.S. National Institute of Standards and Technology launched an AI Agent standards initiative; on August 2, the EU AI Act’s high-risk obligations formally took effect.

The lists are getting longer, but they defend against the same category of things: tools being exploited by external attackers, hijacked by malicious instructions, induced to exceed permissions and call other systems. The design premise of the entire defense line is that the vendor is on the user’s side, and threats come from outside.

ZCode’s and Grok Build’s exfiltration channels happen to sit right in the blind spot of that premise. They’re not on the Agent’s capability list, not governed by permission approval workflows, run outside the tool’s execution loop, and even the AI assistant itself can’t perceive their existence. Compare these behaviors one by one against any existing security framework, and not a single alarm triggers.

There is an even more unsettling layer to this: all of these incidents were discovered by accident. One came to light because a configuration error leaked source code, one because a security researcher actively captured traffic, and one because a blogger noticed something odd about disk space. None of them came from a vendor’s own internal review, an industry audit, or a regulatory inspection.

Some have suggested that Agent data behavior should be audited the way public companies’ financial statements are audited. That analogy is half right. The form—regular, standardized, independent third-party review reports that buyers can understand—is correct. But financial audits examine ledgers that companies are legally required to keep. There is currently no regulation requiring vendors to retain records of “what data left the user’s computer,” so the evidence itself is insufficient. Moreover, Agent clients may update weekly, and some poll remote configurations every hour to change their own behavior. An annual audit report is already outdated the moment it is issued.


III. The Real Structural Problem

Taken apart, the details above are three separate incidents at three companies. Put together, they are the same problem. In the ZCode and Grok Build scenarios, the person clicking the “Agree” button is an individual developer, but the person bearing the consequences of a data breach is their employer and their clients. The latter never appears in any consent flow from start to finish, and has no channel through which to learn that their code was ever packaged and uploaded.

The person who authorizes and the person who bears the risk are not the same person. This misalignment cannot be solved by writing clearer pop-ups or making toggles more prominent. Individual-level informed consent structurally cannot solve this problem. When an employee opens a company project with a personal account, the only thing they can consent to is their own portion of authorization. The company’s portion—they have no standing to consent to it.

For enterprises, this means the questions to ask when procuring an Agent need to change. Not “Is this Agent secure?” but rather: Do the things I authorize and the risks I bear reconcile?

Following this line of thinking, there is one more thing worth spelling out: promises cannot be verified. “Data is destroyed immediately after use,” “not retained,” “encrypted upload”—these phrases answer how long data is retained and whether it can be intercepted by a third party in transit. They cannot answer several other questions: Has the data already left this machine? Who has access to it on the server during processing? Who holds the decryption capability for the encrypted package? What deletion policy was applied to previously uploaded data?

Looking at encryption alone: if the key is on the server side, then “encrypted upload” only proves security in transit—it does not follow that the vendor themselves cannot decrypt it. Whoever holds the key holds the decision-making power. So for enterprises, verifiability is worth more than promises. An architecture that can be externally verified beats ten pages of written commitments.


IV. Six Lines of Defense Enterprises Should Demand from an Agent

Translated into a procurement checklist, the problems above come down to roughly six items.

First: It must be able to live in your own house

This is the only one that solves the problem at its root. All the incidents above share one necessary precondition: data must first leave your network before it can reach someone else’s hands. Private deployment is about cutting off exactly this step. How to tell: cut off the external network egress from the deployment environment and see how much functionality remains. If your Agent cannot even do knowledge base Q&A after being disconnected, that means its data was outside all along. Here you need to distinguish between two kinds of “privatization.” One is where data is stored in your own data center, but every model inference request still goes to the vendor’s cloud. The other is where inference also runs on your internal network. Only the latter truly keeps data within your domain.

Second: It must make every action visible

Audit logs must at minimum answer four questions: who, when, on what resource, did what. If any one of these four is missing, the logs are just decoration. How to tell: pick a call from last Wednesday afternoon and ask your ops team who initiated that session, which tools were invoked along the way, what the parameters were, what was returned, and where the output files were stored. If they can give you a complete chain within five minutes, the logging is usable. If they have to dig through server logs to piece it together, it might as well not exist.

Third: It must know who can touch what

In a permission model, the most important half is often not the “allow” half but the “deny” half. A well-designed permission system should evaluate in this order: explicit allow > explicit deny > inherited from parent > default deny when unset. That last one is the key—it determines whether newly added features are off or on by default. How to tell: find a business user, have them access an application API they don’t have permission for, and see whether it returns a 403 or an empty list. A system that returns an empty list has its permission filtering written into query conditions, and this kind of implementation often has gaps in batch APIs and export APIs.

Fourth: It must act inside a cage

Agents need to execute code, read and write files, and run commands. If these actions happen on the main system, a single misoperation is enough to write up an incident report. How to tell: see where its code execution happens. A reasonable approach is a separate isolated environment per session, reclaimed when the session ends. You can have it run cd / in the sandbox to see what it can see, then delete the session, create a new one, and check whether what you just deleted is really gone.

Fifth: It must control its mouth and its memory

There are two things here. One is content moderation: there should be a sensitive word library, both input and output directions should be screened, and moderation actions should be tiered—block, warn, replace, log only—with different levels of force for different scenarios. Moderation configuration should ideally be isolated by organization; otherwise, different compliance requirements from different departments cannot be satisfied simultaneously.

The other is more easily overlooked: memory. As mentioned in the previous article, model parameters do not change at all after a conversation ends. You think it remembers you because the system puts the conversation history back into the input for the current round. So memory is a piece of data, not a capability. If it is data, it must answer where it is stored, who can see it, and how long it is kept. How to tell: ask how many tiers its memory has, whether sensitivity can be set per memory entry, whether sensitive memories undergo permission checks during retrieval, and whether access records are left when memories are read. A mature memory system should have an out-of-the-box “privacy-first” policy so enterprises don’t have to configure from scratch.

Sixth: It must be replaceable

This one corresponds to the “AI relay station” type of risk, and also to vendor lock-in. Not being bound to a specific large model means you can use one vendor today and switch to a self-built inference service on your internal network tomorrow, without changing a single line on the business side—Agent, knowledge base, or workflow. The way to achieve this is through open protocols, such as the OpenAI-compatible protocol, where any endpoint implementing that protocol can be plugged in. How to tell: ask “How is your model layer connected?” If the answer is the name of a specific model company, then you are buying that model company’s service, not a platform.


V. What These Six Items Look Like on a Single Platform

The six items above are not paper standards—they correspond to concrete engineering practices. Take AgentSteamer as an example, and let’s go through a few points directly relevant to this article’s topic.

  • Where it lives. AgentSteamer supports full private deployment, with core data such as knowledge bases, artifacts, and session records all kept inside the enterprise—data never leaves the domain. At the model layer, AgentSteamer is not bound to a specific large model and connects to any endpoint via the OpenAI-compatible protocol, with OpenAI, Azure, DeepSeek, Qwen, Ollama, and vLLM all within its support scope. This is very practical for enterprises: the inference service can be swapped for a self-built one on the internal network, and the “relay station” link mentioned earlier disappears. The built-in embedding model BAAI/bge-m3 and reranking model BAAI/bge-reranker-v2-m3 are both MIT-licensed and commercially usable, requiring no external API calls.
  • Visibility. AgentSteamer’s audit logs record user, operation, resource, IP, UA, and status, with multi-dimensional filtering by operation, resource, user, and date. AgentOps provides session tracing, so you can see on the interface which nodes a call went through and which skills and MCP tools were invoked.
  • Who can touch what. The permission model is a matrix of 11 resource types by 10 operation types, with resources covering applications, workflows, models, knowledge bases, tools, prompts, users, roles, logs, files, and system configuration. The evaluation order is explicit allow > explicit deny > inherit > default deny. Model provider API Keys are masked, and non-administrators cannot see the plaintext.
  • Acting inside a cage. The platform integrates an independent sandbox service, where each session has its own separate directory, isolated from the main system, and the directory is cleared when the session is deleted. That earlier line about “having it run cd / in the sandbox” is exactly how it is designed in AgentSteamer.
  • Controlling mouth and memory. On content moderation, the sensitive word library supports categories, severity levels, and replacement words; moderation actions are divided into four levels—block, warn, replace, log—covering both input and output directions, with configuration isolated by organization. On memory, AgentSteamer uses a three-tier memory structure, with policies inherited level by level along “node → application → organization → platform.” Configuration items include a default sensitivity value and a permission check toggle, and the platform ships with five default policies, one of which is privacy-first. Every read, write, retrieval, and injection of memory leaves an access record containing user, IP, UA, session, and application information.
  • Replaceability. The way the previous items are implemented all points to the same thing: AgentSteamer is developed by Shanghai Immersivalley Information Technology Co., Ltd., and its model layer, tool layer, and knowledge layer are all replaceable components rather than a rigidly bound whole. The tool side follows the MCP protocol and Agent Skills specification, and skill packages can be imported and exported.

A few other things done along the way: there is a review process before publishing, routed at three levels—department, company, and application—where with multiple concurrent reviewers, any one approval completes it; deleted resources go to a recycle bin first and are only truly purged after the retention period; APIs have rate limiting. These don’t solve the fundamental problem, but they give an answer to the question of “can we trace it after something goes wrong.”


VI. Four Verifications You Can Do Right Now

With the reasoning out of the way, here are a few actions you can verify yourself without waiting for vendor cooperation.

  1. Pull the network cable. Cut off the external network egress of the deployment environment and walk through the core business. This step tests whether “data doesn’t leave the domain” is an architecture or just talk.
  2. Capture packets. Use standard network packet capture tools to see exactly what outbound connections this machine has. cereblab’s analysis of Grok Build was done with ordinary packet capture tools, without any special means. A normal enterprise Agent should only have the few outbound connections you explicitly allow, and you should be able to say what each one is for.
  3. Find the key. If the product has features like “encrypted upload” or “encrypted sync,” ask one question: who holds the key to this package. If the key is on your side, encryption is effective for you; if the key is on their side, encryption is only effective against third parties on the road.
  4. Ask about the authorization chain. Who has the right to consent to data being collected, and who bears the consequences of a data breach — are these the same person? If they are the same, the popup is effective notice; if not, no matter how clearly the popup is written, it’s just going through the motions.

Three of these four don’t require the vendor’s cooperation. This probably also illustrates the nature of the problem: whether data is safe ultimately depends on whether you can see and verify it yourself, not on how the other party guarantees it.


Seven, Frequently Asked Questions (FAQ)

What is the biggest data security risk when enterprises use Agents?

It’s not hacker attacks, it’s the tool itself exfiltrating data. In multiple incidents exposed in 2026, the source of risk was the companies providing the tools themselves: Zhipu ZCode packaged and uploaded projects along with historical modification records, Grok Build transmitted files that users had explicitly declared “do not read,” and Claude Code was reported by the Ministry of Industry and Information Technology’s NVDB for backdoor risks. These channels are not covered by existing security frameworks, because they assume by default that vendors and users are on the same side.

狄, 大人
AgentSteamer Content Team