# Python Integer Caching: Why 256 is 256 but 257 is Not

**Why does `256 is 256` evaluate to `True`, but `257 is 257` return `False`? Python pre-allocates small integer objects (traditionally -5 to 256) inside a global array to save memory. Because this optimization is version-dependent, using the `is` operator for numeric comparisons introduces subtle, breaking bugs.**

Run `a = 256; b = 256; a is b` in your Python REPL, and it returns `True`. Run `x = 257; y = 257; x is y` on separate lines, and it returns `False`. This behavior isn't a glitch—it is a direct consequence of how CPython optimizes memory allocation for low-value integers.

Understanding the mechanics of this behavior requires looking beneath Python's high-level syntax and into CPython's memory management.

## How does CPython optimize small integers under the hood?

CPython pre-allocates a static array of integer objects for values between `-5` and `256` (inclusive) during interpreter initialization. When you assign one of these values, CPython returns a pointer to the pre-existing object rather than allocating new memory on the heap.

In CPython, integers are not lightweight, 8-byte primitive types like they are in C or C++. Instead, they are defined as `PyLongObject` structures. Every `PyLongObject` contains metadata inherited from the base `PyObject` structure, including:

*   `ob_refcnt`: An 8-byte reference counter used for garbage collection.
*   `ob_type`: An 8-byte pointer to the type object (`&PyLong_Type`).
*   `ob_size`: An 8-byte signed integer indicating the size of the digit array.

On a 64-bit system, this structural overhead means even the integer `0` consumes 24 to 28 bytes of memory. Constantly allocating and deallocating these structures for basic loop counters or arithmetic operations would cause massive heap fragmentation and CPU overhead.

To mitigate this, CPython's source code (specifically in `Objects/longobject.c`) defines an array named `small_ints`. During startup, CPython initializes this array with `PyLongObject` instances for every number from `-5` up to `256`. 

## Why does the "is" operator fail on larger numbers?

The `is` operator evaluates object identity by comparing memory addresses directly, whereas `==` evaluates value equality by invoking the object's comparison method. For integers outside the cached range, CPython allocates distinct heap memory addresses, causing `is` to return `False` even if the values are identical.

When you assign a number outside the `small_ints` range (such as `257`) across separate lines in the interactive REPL, the CPython runtime cannot reuse a cached instance. It calls `PyLong_FromLong()`, which invokes `_PyObject_New` to allocate a completely new `PyLongObject` on the heap.

As a result, the variables point to two entirely different addresses in virtual memory. The `is` operator checks if the raw pointer values are equal. Because these pointers point to distinct heap allocations, the identity check fails, even though the underlying numeric values match.

## How does Python 3.15 change integer caching behavior?

Starting in Python 3.15, CPython is raising the small integer caching threshold from `256` up to `1024`. This means that expressions like `1000 is 1000` will evaluate to `True` in Python 3.15, whereas they evaluated to `False` in all prior versions.

This upcoming adjustment highlights why relying on the `is` operator for value comparison is a dangerous anti-pattern. The boundaries of the `small_ints` array are internal implementation details, not part of the Python language specification. 

If you write code that accidentally relies on this cache, your application might pass its tests locally on a newer Python interpreter but fail silently when deployed to a production environment running an older version of CPython.

| Python Version | `256 is 256` | `257 is 257` | `1024 is 1024` | Safe Comparison Operator |
| :--- | :--- | :--- | :--- | :--- |
| **Python 3.12 and older** | `True` | `False` | `False` | `==` |
| **Python 3.15+** | `True` | `True` | `True` | `==` |

## FAQ

### Why does CPython's cache start exactly at -5?
CPython starts the cache at `-5` because low-value negative numbers are frequently used as index offsets (such as accessing lists from the end with `-1`) and error-state sentinel values in C-API integrations.

### Does CPython optimize string objects using a similar mechanism?
Yes, CPython uses a process called "string interning" for short, ASCII-only strings that look like Python identifiers. This is handled via the `sys.intern()` function and internal dictionary keys, allowing the interpreter to perform fast pointer comparisons instead of expensive character-by-character string comparisons.

### Why do two large identical numbers sometimes return True for "is" inside a script?
When Python compiles a single script file or module, the compiler processes entire code blocks at once. During this compilation phase, CPython performs constant folding and merges identical literals within the same code object's constant pool (`co_consts`), causing them to share the same memory allocation.
