본문으로 건너뛰기
L10.8

Code Execution as a Tool

Goal

Explain why code execution is a high-authority tool and design a bounded execution envelope with limits on code, files, network, time, processes, and returned output.

Code can calculate things that are awkward for a language model to do reliably. It can also read files, open network connections, start processes, consume resources, and change data.

That makes “run code” fundamentally different from a narrow calculator function.

Start with the minimum capability​

If the task only needs arithmetic, expose arithmetic. If it needs dataframe operations on one uploaded file, give the executor only that file and the needed libraries.

A general-purpose interpreter should be the result of a deliberate requirement, not the default tool.

A sandbox limits blast radius​

A code-execution boundary can restrict:

  • filesystem paths;
  • network access;
  • environment variables and secrets;
  • allowed packages or binaries;
  • CPU time;
  • memory;
  • process count;
  • output size;
  • writable locations.

No sandbox is made safe by prompt wording. The restrictions must come from the runtime or surrounding application.

Inputs and outputs need boundaries too​

Do not interpolate untrusted text into shell commands when a structured library call would work.

Likewise, large stdout or generated files should be capped and inspected before being placed back into model context. A program that prints a million lines can exhaust context even if it never escapes the sandbox.

Timeouts are not only convenience​

An accidental infinite loop can consume resources. A timeout converts that failure into a bounded, observable result.

The workflow should distinguish:

completed
runtime_error
timeout
resource_limit
policy_blocked

These categories guide different next steps.

Generated code still needs review for high-impact actions​

Even inside a sandbox, some tasks can have meaningful side effects—for example modifying an approved working file that will later be published.

Use human approval or deterministic checks where the consequence justifies it. “The model wrote valid Python” is not the same as “the action is acceptable.”

Distinguish code generation from code permission​

A model can propose code that is logically correct for a calculation and still request capabilities the current job should not have. Review therefore happens on two axes: what the program computes and what the runtime lets it touch. Sandboxing answers the second question even when the first answer is uncertain.

For example, a CSV analysis may legitimately read /workspace/input.csv and write /workspace/result.csv. It does not need to read ~/.ssh, inspect environment secrets, or make outbound requests. A narrow mount and denied network make those boundaries real regardless of generated code content.

Reproducibility needs environment identity​

If code output matters to a decision, record the interpreter/runtime version, allowed packages, input artifact hashes or IDs, resource limits, and generated source or source hash. “It ran successfully” is weak evidence if another person cannot recreate the same execution envelope.

Predict

A task only needs to sum a list of numbers. Which tool design has the smallest useful authority?

Run the local Lab​

python labs/notebooks/level-10/l10-08-code-policy.py

The Lab checks three code-execution requests against a policy with no network, a 3-second limit, a 2,000-byte output limit, and one writable folder.

  1. Run the command. You should see calc -> allowed, web -> policy_blocked_network, and large-output -> resource_limit_output.
  2. Notice the order of checks in check: network permission first, then time, then output size, then write paths. The first failing rule is the one reported.
  3. Lower the output limit: change "max_output_bytes": 2000, to "max_output_bytes": 50, and rerun.
  4. Now even calc, which prints only 100 bytes, becomes resource_limit_output. A limit that is too tight blocks legitimate work, so limits must fit the task. The script then stops with an AssertionError, which is expected: its final checks were written for the original values. Change the value back afterward.

Loading lab…

Quick Check

1. Why is a general code-execution tool high authority?
2. What should enforce network denial for a sandboxed job?
3. Why cap returned stdout?

0 of 3 questions answered.

Key Takeaways

  • General code execution has much broader authority than narrow tools.
  • Sandboxing must enforce filesystem, network, resource, and secret boundaries.
  • Prefer structured library calls over shell interpolation.
  • Time and output limits turn runaway behavior into bounded failures.
  • High-impact generated actions may still need deterministic checks or approval.

Next Lesson

Next, move from text-only evidence to images and learn how visual inputs add another source of observations and ambiguity.

References

Lesson actions

Completion is stored locally on this device.

View progress