1. Helpers
  2. cache

New to Gradio? Start here: Getting Started

See the Release History

To install Gradio from main, run the following command:

pip install https://gradio-builds.s3.amazonaws.com/15ce43720d1611f44e4661c74d604eabe0d04daf/gradio-6.28.0-py3-none-any.whl

*Note: Setting share=True in launch() will not work.

cache

@gradio.cache(···)

Description

Decorator that auto-caches function results based on content-hashed inputs. Works with sync/async functions and sync/async generators. For generators, all yielded values are cached and replayed on hit. Cache hits bypass the Gradio queue. It can also be called at runtime as gr.cache(fn)(*args) to cache intermediate helper calls.

Example Usage

import gradio as gr

@gr.cache
def classify(image):
    return model.predict(image)

@gr.cache(max_size=256, per_session=True)
def generate(prompt):
    return llm(prompt)

Initialization

Parameters
🔗
fn: Callable | None
default = None

The function to cache. When used as @gr.cache without parentheses, this is the decorated function. When used as @gr.cache(...), this is None. When used as `gr.cache(fn)(...)`, this must be a callable.

🔗
key: Callable | None
default = None

Optional function that receives the kwargs dict and returns a hashable cache key, e.g. to only cache based on the prompt, pass in: lambda kw: kw["prompt"]

🔗
max_size: int
default = 128

Maximum number of cache entries. Least-recently-used entries are evicted when full. Set to 0 for unlimited. Default: 128.

🔗
max_memory: str | int | None
default = None

Maximum total memory usage before eviction. Accepts strings like "512mb", "2gb" or integer bytes. When exceeded, least-recently-used entries are evicted. If None, no memory limit is applied. If both max_size and max_memory are set, the cache will evict entries when either limit is reached.

🔗
per_session: bool
default = False

When True, each user session gets an isolated cache namespace, preventing cached results from leaking between users. Per-session entries are cleared when the client session disconnects. The max_size and max_memory limits apply to the sum of all entries across all sessions.