Before handing a coding agent a task, make a short record of the session that will run it: the application, working directory, active sandbox state, and tool route. A checkbox saved yesterday is weak evidence for a different session today.
GitHub announced general availability of local sandboxing on October 7, 2026 for Copilot CLI, the Copilot app, and VS Code sessions using Agent Host. The announcement separates model execution from tool isolation: choosing a model does not establish the permissions of its tools.
This documentation-based guide, checked October 8, offers an original session-review worksheet. It is not a penetration test, a claim that a particular machine is protected, or a report of hands-on Copilot testing. Start with harmless project material you are authorized to use.

1. Identify the surface and the current session
GitHub's sandbox overview says local sandboxing is off by default and that CLI and app settings are separate. It describes local containment as process-level restrictions, not a separate virtual machine. On an unsupported host, a CLI saved preference can remain enabled while the current session runs without a sandbox; managed policies can instead require execution to stop.
Give each review a narrow identity. “Copilot on my laptop” is not enough. Record the surface, version, operating-system version, session, project, and current directory. If you switch from the CLI to the app, make a new record instead of copying the old result.

2. Inspect status before interpreting settings
For the CLI, GitHub documents /sandbox status as the current-session check. Use /sandbox policy to inspect effective access for the current directory. Adding a proposed command, such as /sandbox policy npm install, previews its policy without running that command. Read the report's Notes too.
Enabling sandboxing in one CLI session does not immediately enable other already-open sessions. Check the session where the work will actually happen. Keep the status observation and the policy observation in separate fields, with the time you checked them.
For the app, the configuration guide distinguishes the configured policy summary from a running sandbox. Project defaults affect new sessions; filesystem, network, and credential changes take effect in new or restarted sessions. The app checks host support when the first sandboxed shell starts. An unsupported host or policy can therefore produce a failure after settings were accepted.
Our review rule: if the available evidence is only a saved preference, mark the session unverified. If a warning says the required boundary cannot be enforced, stop the task and resolve that requirement with the responsible administrator. Do not turn a warning into a successful check.
3. Name the tool route
The filesystem-policy guide makes an important distinction: CLI built-in file tools run inside the CLI process and honor policy through software checks, rather than through the operating-system sandbox used for subprocesses. Local MCP and language servers are sandboxed by default; remote MCP servers are not locally sandboxed. Network policy can still restrict connections from the CLI when MCP sandboxing is enabled.
Use that distinction to describe the next operation precisely. “Read a file” could mean a built-in file tool, a shell command, or a remote service. A status result for the local shell is not evidence about a remote service's permissions or data handling.

| Planned route | Record before the task | Do not infer |
|---|---|---|
| Local shell command | Session status, directory, effective policy and warnings | That every other tool follows this route |
| Built-in CLI file tool | Tool identity and the policy it is expected to honor | That an OS sandbox intercepts its file operations |
| Local MCP or language server | Which process is local and its sandbox setting | That the default was never changed |
| Remote MCP service | Service identity, destination and separately approved data scope | That a local sandbox confines the remote server |
4. Copy this session-review card
The following is an editorial template, not a Copilot export or an official security checklist. Record observations without copying tokens, private file contents, or sensitive paths into a shared report.
- Task and allowed material: What small task will run, and what may it use?
- Session identity: Surface, version, OS, execution-host alias, project alias, directory alias, time checked.
- Active state: What does the current session report? Record warnings separately.
- Execution route: Shell, built-in file tool, local server, or remote service?
- Files: Which relevant locations are readable, writable, or denied?
- Network: What internet and local-network access is requested or reported?
- Credentials: Are authenticated Git or GitHub CLI operations available? Record the setting, never the credential.
- Unresolved evidence: What has not been checked, and who can resolve it?
- Decision: Ready for this bounded task, unverified, or blocked. State why.

Do not fill the network and credential rows with “none” simply because sandboxing is on. The configuration documentation lists authenticated Git and GitHub CLI access as enabled by default. It also gives the app and CLI different local-network defaults. Available restrictions vary by operating system, so a policy copied from another machine needs its own review.
5. Use a small synthetic handoff
Example: a teammate asks an agent to summarize a public README. Your record says the intended route is a local shell in the CLI, but the only evidence supplied is an app project-settings screenshot. Mark the task unverified. Ask for the current CLI session status and policy, or agree on a different clearly identified route. This is a fictional review example, not a measured product result.
For another example, suppose a harmless project task stops at an access error. Save the command description and the relevant warning, then compare the required resource with the intended scope. A failed command is useful evidence; making it succeed with broader access is a different decision. Keep that decision outside this read-only review.

Refresh the card after a change of surface, session, directory, policy, host, or tool route. Leave unobserved results blank. The useful handoff is a dated description of this task's boundary and remaining questions, not a general “sandbox tested” badge.