blob: 1f4370dbde0985c1761b115f7389ea8de58606e7 [file] [view]
# Security Policy
This document describes the security model, threat boundaries, and security
recommendations for Protobuf. It is intended to help developers understand what
security guarantees Protobuf provides, and how to safely parse and serialize
messages.
--------------------------------------------------------------------------------
## Reporting a Vulnerability
If you believe you have discovered a security vulnerability in Protobuf, please
report it via
[Google's official vulnerability disclosure channel](https://bughunters.google.com/report).
For reports that do not meet the criteria or bar for the Google Vulnerability
Reward Program (VRP) but are still security-sensitive, you can privately report
them by opening a
[draft GitHub Security Advisory](https://github.com/protocolbuffers/protobuf/security/advisories/new).
If an issue relates to something which is listed as Best Effort or Out of Threat
Model below, please open a
[public issue](https://github.com/protocolbuffers/protobuf/issues) instead.
--------------------------------------------------------------------------------
## Best Practices
* Handle `.proto` files as "code" by default, similarly to the handling of
`.java` or `.cc` files. Executing untrusted code can potentially be a supply
chain attack vector.
* Prefer to use the binary wire format encoding where possible: all
Google-maintained Protobuf implementations treat binary wire format parsing
as the primary use case to handle untrusted inputs. Other formats are
considered secondary.
* **Apply Defensive limits:** Prefer to have a layer that applies defensive
limits or custom filtering around the Protobuf parse: for example gRPC
enforces a 4 MiB payload limit by default. We recommend having sensible
limits defensively applied on all untrusted inputs.
* While all Google maintained Protobuf implementations are considered to be
intended to be used with untrusted inputs, not all runtimes are equally
hardened:
* **JavaProto** (including the Kotlin bindings) is recommended as the
default implementation for the best security posture, as it benefits
from the JVM's memory safety guarantees.
* **C++Proto** is recommended for applications requiring optimal
performance, and has been extensively fuzzed and hardened such that
Google uses it without sandboxing to parse untrusted binary format
inputs in critical surfaces. When using C++, developers must remain
aware of native memory management risks.
* In languages that have multiple supported implementations, strongly
prefer to use the default implementation for best security. The
non-default implementations are generally supported for more exotic
use-cases and may not have as much hardening attention. The default
Python and PHP runtimes use a C extension that has been hardened more
than the fallback behaviors that that do not use any C extension.
* **Use latest releases:** Ensure that the generated code and the runtime
library version match exactly and are kept up to date with the latest patch
release. Certain obscure and low severity issues may only be patched on the
latest release.
## Threat Model and Security Boundaries
Protobuf is a serialization format designed to parse data efficiently. Depending
on the input format, language runtime, and integration context, different parts
of the Protobuf ecosystem are hardened against adversarial inputs.
Our threat model divides surfaces into three tiers:
### Proactively Hardened
These surfaces are fully supported and hardened against adversarial inputs.
* Security vulnerabilities identified in these paths are treated with high
priority, actively fixed, and issued CVEs.
* Where necessary, certain patches that are technically breaking may be done
if it is inherently necessary to close a security issue. This also includes
that gencode-runtime version compatibility guarantees may be broken if it is
strictly necessary to close a security issue (see documentation
[here](https://protobuf.dev/support/cross-version-runtime-guarantee/#exception)).
These cases are rare and we will avoid this wherever it is possible to close
the security concern without a breaking change.
* Google trusts these surfaces enough to use in sensitive, publicly exposed
endpoints without sandboxing.
### Reactively hardened
The Protobuf team welcomes bug reports and pull requests to improve the
hardening of these surfaces. However, we typically will not break backwards
compatibility guarantees to address security issues in these areas (especially
for lower severity risks).
Defensive hardening is applied to these surfaces, but with weaker guarantees
compared to our hardened surfaces.
* Serious security issues are still highly prioritized for fixing on these
surfaces.
* We do not break backwards compatibility guarantees to address lower severity
issues on these surfaces outside of major version bumps. In some cases this
means low-impact known issues may even be left open if they inherently
cannot be closed without a breaking change. Serious issues (like RCE) would
still be urgently addressed.
* When using these surfaces on potentially malicious inputs, especially
security sensitive usages are recommended to consider application-level
isolation (such as sandboxing) or other defensive handling.
### Outside of CVE Threat Model (Best-Effort Hardened)
These are surfaces where Protobuf libraries are not as hardened against
adversarial inputs.
Outside of threat model does not mean we do not care at all about
security-relevant behaviors on these surfaces: we still apply defensive
hardening where feasible, and welcome reports on these surfaces. However, as
these surfaces are expected to be used with trusted inputs, the higher priority
for those cases is other topics including stability, performance, and developer
ergonomics.
## Proactively Hardened
### Parsing of Binary Wire Format Encoded Data **(Primary Use Case)**
This is considered the primary surface of security concern in Protobuf
libraries.
In this use case, the `.proto` schema files, the compiler (`protoc`), and the
generated code are fully trusted. The incoming binary wire format bytes are
untrusted and may be adversarial.
The parsing library will safely process or reject any arbitrary byte stream
without exposing the server to memory corruption, out-of-bounds reads/writes, or
remote code execution (RCE).
Note that the intended surface here is only parsing: once a message is in-memory
it is treated as a trusted object in our threat model. For that reason,
serialization or any other in-memory handling of parsed objects is not
considered a surface within the threat model. For example, parsing an untrusted
wire format payload should not be able to reach uncontrolled recursion, but
serializing an arbitrarily in-memory object may. This is similar in nature to
how modern browsers do not throw `RangeError` on `JSON.parse()`, but do on
`JSON.stringify()`.
### Parsing of ProtoJSON Format
[ProtoJSON](https://protobuf.dev/programming-guides/json/) allows using `.proto`
schemas with standard JSON encoding. This is considered an ancillary supported
encoding, and the binary format should be preferred where feasible.
Under the same threat model as publicly exposed binary wire format services,
ProtoJSON parsing is intended to be used with untrusted inputs. ProtoJSON
serialization is similarly not considered within the threat model risks.
--------------------------------------------------------------------------------
## Reactively Hardened
These areas are defensive against malicious inputs, but application-level
defense (such as sandboxing) is recommended.
### Parsing Text Format
Text Format is designed for local debugging, testing, and managing trusted
configurations by developers. It is not intended to be used as an interchange
format and is not recommended to expose public services which consume Text
Format as an encoding.
Note: Text Format parsing currently does not enforce any depth limit in several
supported runtimes (they support opt-in depth limits). We cannot enforce a depth
limit by default without breaking backwards compatibility, but may begin to
enforce depth limits by default as part of a future breaking change release.
### Lite Runtimes (C++ Lite, Java Lite) Denial of Service Risks
Lite runtimes target mobile and web usage: they are optimized for those
constrained envirnoments, and prioritize small binary size at the expense of
other properties.
Lite runtimes are still intended to be used to parse untrusted inputs, but DoS
issues are considered to be much less severe in mobile contexts, as the
opportunity and impact of an attacker successfully freezing one app is low
compared to reducing the availability of a public service.
We still intend to mitigate such risks, but Lite gencode/runtimes are not
recommended for servers exposing public endpoints.
### 'Wrong Kind of Exception Thrown'
In memory-safe runtimes (Java, Go, Python, C#), reaching a catchable exception
or runtime error (such as `IndexOutOfBoundsException` or `StackOverflowError`)
instead of the declared exception (such as `InvalidProtocolBufferException`) is
treated as an important bug to fix, but is considered to be a minor security
concern relative to serious issues like native heap corruption, data leakage,
remote code execution, or DoS vectors from unbounded computation or memory use.
It is recommended that RPC handler code be defensive against unexpected
exceptions if they are exceptionally sensitive to such behavior.
## Outside of CVE Threat Model (Best-effort Hardened)
### `protoc` CLI
The `protoc` CLI is an offline developer tool. `.proto` files are considered
source code (equivalent to `.java` or `.cpp` source files).
A supported modality of the CLI is to parse untrusted `proto` files to emit
`FileDescriptors` which can enable further machine processing of the schemas.
While `protoc` is hardened on a best-effort basis for this use case, we
recommend using defensive validation and sandboxing whenever running `protoc`
against potentially malicious inputs.
Untrusted flags being passed to `protoc` is fully outside of our threat model:
CLI flags are never be adversarial and arbitrary behavior driven by CLI flags
may be working as intended.
Caution: Compiling and executing generated code from untrusted schemas is
functionally equivalent to compiling and running arbitrary third-party `.java`
or `.cpp` source code and executing it: there may be intentional language
features which act as intentional code injection into the generated code. You
must treat untrusted `.proto` files the same as any other programming language
code in terms of supply-chain risk in this way, and not blindly execute the
gencode which was generated off of untrusted `.proto` files.
### DynamicMessage on Untrusted Descriptors
Protobuf supports encoding schemas into a Protobuf message format (e.g.
`FileDescriptorSet` or `DescriptorProto`). These messages can be handled as any
other Protobuf type. Parsing untrusted binary-encoded DescriptorProto falls
within the "primary use-case" described above.
In addition to simply processing DescriptorProto, it is additionally possible in
most runtimes to use a type named `DynamicMessage` which allows for using
runtime-loaded descriptors instead of using generated code and to use that type
with the reflection APIs.
For use-cases sensetive to DoS risks, it is recommended to use `DynamicMessage`
only with trusted descriptors (via trusted side channel source / config pushes).
When using `DynamicMessage` with a descriptor sourced from an untrusted source,
you may need to validate and sanitize them as you would user provided SQL.
Caution: Usage of `DynamicMessage` with malicious descriptors reaching an RCE or
information leak would still be treated as a high priority issue. However, there
are inherently reachable cases of where malicious descriptors used with
`DynamicMessage` can reach behavior which may otherwise be considered a Denial
of Service risk under our primary threat model. For example, it will be
reachable to hit memory use which is O(N*M) where N is "# of messages observed
on the wire" and M is "size of the message definition". Since untrusted
descriptors gives an affordance for arbitrarily large message definitions, using
DynamicMessage with untrusted descriptors and untrusted binary format inherently
can have memory amplification risks.
### Adversarial Application Code
Violating runtime API constraints or passing invalid arguments directly to a C++
API is considered an application integration error rather than a library
vulnerability.
Protobuf libraries do harden against the impact of certain classes of mistakes
being worse; for example, we often will panic if we can detect that an out of
bounds memory reads will occur in some cases. This is considered
defense-in-depth and misuse is not considered a vulnerability.
Excepting the surfaces enumerated above as hardened, Protobuf APIs in
memory-safe languages reaching memory safety problems on 'bad' parameters is
considered an high priority bug, but typically not within scope for CVE
disclosure.
In languages like PHP, this means that use of PHP Protobuf in an unconstrained
multi-tenant system where malicious application code may try to attack other
jobs concurrently running is not within our threat model.
Examples:
* Passing negative or invalid buffer size parameters directly to
`ParseFromArray` in C++ is wrong application code. It may be hardened to
panic instead of risk out of bounds memory reads, but is not intended to be
gracefully handled as a malformed-wire-bytes input would be (following C++
idioms).
* In a memory language Python, if code like `msg.repeatedField[-2147483649]`
can reach a segfault, that is considered an important bug to fix, but it is
not considered to be within CVE scope.
### Differential Parsing (Gateway propagation of original payload)
Differential parsing is a risk stemming from by two different libraries parsing
the same data with different interpretations.
In some contexts and for some formats differential parsing is considered a
security-sensitive topic. The primary risk is around flows that would validate
in a gateway, forwarding the data unmodified, and then a backend handles the
original payload and interprets the data differently, bypassing the intended
checks.
The Protobuf binary format is explicitly designed for propagation of unknown
fields, where the gateway may not be aware of the content at all and forwarding
will result in the next server who has an updated version of the schema will
corresponding parse to a different interpretation because it is aware of those
fields.
Additionally, there may be certain edge-case byte sequences where a few
different interpretations which may be considered acceptable within spec. Google
maintained runtimes will never encode these sequences, but they may successfully
parse them.
When using ProtoJSON format, the underlying JSON format itself contains
significant inherent ambiguities (as noted in ECMA-404 and RFC-8259, including
that there is no spec around the handling of duplicate keys and numeric
precision). As ProtoJSON is built on top of that foundation, ProtoJSON inherits
these ambiguities where certain sequences may have multiple different spec
permissible implementations. Spec Protobuf implementations strongly attempt to
avoid ever encoding such sequences (including that they always quote large
int64s, and don't emit duplicate fields), but the parsing behavior may differ in
such sequences and ProtoJSON cannot spec behavior which is unimplementable when
using ecosystem of JSON parsers.
For both formats, an architecture where a gateway performs validation and
forwards the original user request to a second server which reparses the user's
request but presumes validation has already occurred is outside of our threat
model.
For best security practice, it is recommended to:
* Use a different set of messages schema for your public API and internal
messages. Besides the security benefits, this decoupling also allows for
easier evolution of your system, where public APIs often need to change
slowly but internal ones can evolve faster.
* Where you do propagate the same message type, always prefer for the gateway
to parse and reserialize instead of forwarding the original payload
verbatim, as this will commonly normalize edge case byte sequences.
* Wherever possible to have each microservice verify any relevant ACLs for
actions it is taking based on its interpretation of the request rather than
rely on gateway validation.
### Specific depth cap exceeded but without uncontrolled recursion
Protobuf parsers apply depth limits (which are configurable): the purpose of
these depth limits is to prevent resource exhaustion issues stemming from
uncontrolled recursion.
To generally maintain consistent and interoperable behavior, we intend these
depth limits to be consistent in behavior in what payloads will be accepted or
rejected for a given integer depth.
Issues where an edge case is successfully parsed which is deeper than the exact
intended limit, this is viewed as a simple bug as long as it does not expose
meaningful resource exhaustion risks.
We welcome reports and patches for issues of that nature, but do not view it as
a security concern and so these issues can be filed via our public GitHub Issues
flow.
### Canonical Representation and Signature Verification
There is **no canonical representation** of Protobuf messages.
* **Deterministic Serialization:** Many runtimes support deterministic
serialization, which guarantees that a given build of a binary will
serialize the same message to the same sequence of bytes. However,
deterministic does not mean canonical: rebuilding the binary, changing the
compiler version, or minor schema modifications can legally result in an
alternate serialized byte representation that would have the same
interpretation when parsed. For more details, see
[Protobuf Serialization is Not Canonical](https://protobuf.dev/programming-guides/serialization-not-canonical/).
* **Recommendation:** Do not use the serialized byte output of Protobuf
messages to compute stable cryptographic signatures. You may still sign a
given encoded byte sequence, knowing that there are other byte sequences
that would be equally valid representations of the same message. If you need
a stable signature of a given message, you must define and implement your
own canonicalization specification over the parsed message fields and not
over the encoded messages.
### Risks if wire bytes are modified in-transit
In terms of transport security, Protobuf is functionally equivalent to a
plaintext format: the encoding has no built-in signing or other integrity
features.
Best practice is to transport Protobuf encoded data over https. If signing or
other integrity features are needed, it is expected to be done in the layers
built top of the Protobuf libraries.
### API surfaces which not intended for direct public use
Protobuf has APIs which are not advertised or intended for direct public use.
These APIs may have non-obvious invariants for how they must be used.
Most notably, `upb` is a library which is used as an implementation detail API
of our other Protobuf libraries to use. `upb` itself is a highly optimized C
library which requires callers maintain invariants to be sound.
Security issues may arise if our language-specific runtimes which use `upb` do
not maintain those necessary invariants, or if `upb` has reachable bad behavior
when all intended invariants are maintained. However, it is not considered not a
security topic if arbitrary bad behavior may be reachable if `upb` APIs are
directly misused (including that `upb's` APIs accept MiniDescriptors/MiniTables
which are considered trusted types, and so will have arbitrary behavior if those
types do not meet the intended invariants).