On July 29, 2026, the GCC Steering Committee adopted a policy rejecting legally significant contributions that contain code generated by a large language model, or LLM. The restriction also covers contributions derived from model output, even after human revision.

That puts the GNU Compiler Collection on a different path from LLVM, which permits AI-assisted contributions when contributors disclose substantial assistance and remain responsible for the code. Both projects maintain compilers, so the difference deserves more attention than a simple comparison between projects that allow AI and projects that don't.

GCC compiles the Linux kernel, Python, GCC itself, and roughly half the foundational software running on servers. Its contribution rules affect a toolchain used far beyond its own developer community. The new policy reflects concerns about copyright ownership, code provenance, and the difficulty of reviewing changes that can affect the behavior of compiled programs.

Where GCC draws the line

The policy rejects “legally significant contributions which include LLM-generated content or are derived from LLM-generated content.” The restriction on derived material is important. Generating a patch with Copilot or ChatGPT and then revising it extensively doesn't make it eligible if the submission remains substantially based on that output.

The threshold for legal significance follows the Free Software Foundation's longstanding guideline of roughly 15 lines or more. This is the same threshold used for GCC copyright assignments. Significant contributions require a copyright assignment or disclaimer from the contributor, and AI-generated content raises questions about what rights the contributor can assign.

The policy leaves several uses of AI available:

  • Developers can discuss ideas with an LLM, research unfamiliar areas, ask for explanations of existing code, and investigate potential bugs.
  • There is an explicit exception for AI-generated test cases, such as small programs that reproduce a bug or check a particular behavior.
  • Trivial contributions, including one-line typo fixes and obviously mechanical changes, can come from AI if that use is properly disclosed and the changes meet normal contribution standards.

The distinction is between using an LLM to help understand a problem and submitting substantial implementation work produced by one. Contributors still need to pay attention to the exceptions, but extensive editing alone doesn't remove a patch from the restriction.

Copyright and correctness create separate concerns

The FSF has required copyright assignments for significant GCC contributions since the 1980s. Consolidating copyright gives it a way to enforce the GPL without having to locate every contributor whose work appears in a disputed piece of code. That legal structure supports GCC's continued distribution under the GPL and limits attempts to turn covered code into proprietary software.

AI-generated code complicates that arrangement. As of August 2, 2026, the legal treatment of model output remains unsettled. Questions being worked through in the US and EU include whether outputs are derivative works of training material, whether anyone holds copyright in them, and what obligations might follow. Possible rights holders include the user or vendor, while some output may have no copyright protection at all.

For GCC, uncertainty about ownership affects an existing enforcement strategy. A contributor may be unable to document the provenance of generated material or establish which rights can be transferred to the FSF. Accepting that material could introduce a legal uncertainty that ordinary copyright-assignment procedures aren't designed to resolve.

Correctness is a separate reason for caution. A compiler bug can affect the programs built with it. Incorrect register allocation under particular optimization flags, for example, can produce a binary that behaves incorrectly despite valid source code. An error in constant folding, where the compiler calculates a value during compilation, can have similarly quiet consequences.

These failures can be difficult to trace because the application code may look correct. Compiler developers with decades of experience have introduced bugs that survived for years. Human authorship is no guarantee of correctness, but a contributor who understands a change and can explain its assumptions gives reviewers something more useful than plausible-looking code alone.

LLVM has chosen a different responsibility model

LLVM's “human in the loop” policy allows AI-assisted contributions. Contributors must be able to answer questions about their code during review, and substantial AI assistance must be disclosed in the pull request.

The Linux kernel, Mesa, Firefox, and Ghostty have taken similarly permissive positions with disclosure requirements. Debian is still debating whether to ban AI-generated code entirely, while the Rust project has moved to restrict LLM use in contributions.

Different policies can reflect different legal structures, review practices, and relationships with contributors. A disclosure rule that suits an application framework may involve different tradeoffs for a compiler used to build safety-critical software. Even two compiler projects can reasonably place different weight on copyright assignments and contributor accountability.

By mid-2026, governing AI contributions has become a practical issue for infrastructure projects. Projects without a policy are likely to face the question through incoming patches, whether or not maintainers have set aside time to decide it.

The rule is clearer than the enforcement

GCC's policy doesn't come with a reliable technical way to identify violations. Automated detectors of AI-generated code aren't dependable enough to make consequential acceptance decisions. A developer could use Copilot to produce 80% of a patch, revise it, and claim human authorship. The diff alone wouldn't prove otherwise.

That limits enforcement at scale, but it doesn't make a written rule useless. Policies also work through contributor expectations and deterrence. GCC's established contributors have reasons to protect the project and follow its rules without a detector checking every submission.

Ordinary code review will still catch some problematic contributions. Poor explanations, unexplained design choices, and a failure to respond to review can expose weaknesses regardless of which tool produced the code. Those signals help maintainers judge a patch, though they don't establish its provenance.

The more difficult case is a capable, well-intentioned contributor who doesn't know whether revising generated code makes it acceptable. The explicit restriction on derived material answers that question. Clear guidance can change behavior even when violations are hard to prove.

Why the test-case exception is useful

The exception for generated test cases is one of the more useful parts of the policy. A compiler test is usually a small program designed to reproduce a particular bug or verify expected behavior. That gives reviewers a narrower question to investigate than a change to the compiler's implementation.

Writing a useful test doesn't necessarily require expertise in register allocation or intermediate representation lowering, the process of converting a compiler's internal representation into a form closer to machine code. It does require understanding the behavior being tested. An LLM can help generate small, focused programs aimed at edge cases, and those programs can then be checked against the intended result.

A test that fails to reproduce the target bug has missed its purpose in a relatively direct way. A flawed implementation patch may appear correct while introducing failures that emerge only under particular combinations of source code, target architecture, and optimization settings. That difference helps explain why the committee would permit generated tests while rejecting substantial generated implementation code.

It doesn't, by itself, settle the copyright questions around generated tests. It does make the exception operationally understandable: the contribution has a limited purpose and a result that reviewers can check. The policy leaves room for a useful application of AI without accepting it throughout the codebase.

What contributors and maintainers need to decide

For significant GCC implementation patches, contributors need to write the code without basing it on LLM-generated material. They can still use models for the permitted research and explanation tasks. They also need to understand and defend the submitted code, as they would under normal review.

Other infrastructure maintainers don't have to adopt GCC's approach. LLVM's disclosure model is also defensible: it allows the tool while keeping responsibility with the contributor. The practical requirement is to document a position before a major submission creates a dispute. As AI-assisted contributions accumulate, deciding how to treat them retrospectively becomes harder.

The GCC Steering Committee has said it will revisit the policy in early 2027. By then, case law and vendor policies may have clarified some copyright questions. LLVM's disclosure-based approach may also prove more sustainable in day-to-day operation than a prohibition that depends heavily on contributors reporting their own tool use.

For now, GCC's restriction is a defensible conservative choice for a project with a copyright-assignment system and demanding correctness requirements. The early-2027 review gives the committee a chance to reconsider that choice as the legal position and experience with other projects' policies develop.