Python Website Requests Sometimes Become Much Slower After the Application Has Been Running for a While

Hi All,

I am experiencing one specific performance problem with my Python-based website where requests to one of my application endpoints sometimes become significantly slower after the website has been running for a while. The endpoint works normally when the application is first started, and typical requests usually complete within the expected amount of time, but after the application has been handling traffic for some period, the same type of request can suddenly take much longer to return. The slowdown is intermittent rather than affecting every request, and restarting the Python application temporarily brings response times back to normal. I am trying to understand what could cause a Python web application to gradually develop slower request processing without an obvious application error or complete failure. The endpoint itself continues returning the correct response when it is slow, so the issue is not that the request crashes or produces incorrect output; the problem is specifically that the response time increases unexpectedly after the application has been running and serving requests.

The affected endpoint performs normal application work and accesses data needed to construct the response, so I have been comparing requests made shortly after the application starts with requests made later in the same process. When the application is freshly started, the endpoint responds consistently and the website feels normal. After the application has processed more requests, however, some requests can take considerably longer even though they are handling similar input and returning roughly the same amount of data. Other requests made around the same time can still complete quickly, which makes the behavior difficult to reproduce from a single request. The application does not show an obvious Python exception when the slowdown occurs, and the process remains available to accept additional requests. Because restarting the application temporarily clears the slowdown, I am wondering whether something inside the long-running Python process is accumulating over time and causing certain requests to spend more time waiting for resources or performing work than they do immediately after startup.

I have already tried measuring the response times instead of relying only on how the website feels in the browser. I have compared requests during a fresh application session with requests made after the application has been running for a longer period, and the difference is visible in the server-side timing as well. I have also checked the application logs for obvious exceptions around the time of the slow requests, but there is no consistent error that appears to explain why the response becomes slower. I have looked at the endpoint’s normal execution path and confirmed that the same general code is being executed for both fast and slow requests. I am now trying to determine whether the delay is occurring inside Python code itself or while the application is waiting on another resource. Since restarting the process temporarily restores normal performance, I suspect that the long-running state of the application may be relevant, but I do not want to assume that this is caused by memory usage or another specific resource without being able to measure it properly.

One thing I would particularly appreciate guidance on is the best way to diagnose this type of gradual slowdown in a Python web application. I want to be able to identify exactly which part of a slow request is consuming the additional time rather than simply restarting the process whenever the problem appears. For example, I am interested in knowing whether Python’s profiling tools can be used safely to compare a normal request with a slow request in a running web application, or whether there are recommended approaches for collecting timing information around individual functions, database operations, and external calls. I am also wondering whether there are common application-level causes of this pattern, such as objects or data structures that unintentionally grow over time, connections that are not released promptly, or background work that gradually affects request handling. I am not looking for a list of unrelated possible problems so much as a recommended diagnostic process that can show where the additional request time is actually being spent.

The issue is becoming difficult to work around because restarting the Python application is not something I want to rely on as part of normal website operation. While a restart temporarily returns the endpoint to its normal response time, it does not explain why the slowdown develops in the first place, and it also interrupts the running application. I would prefer to identify whether there is some resource or state within the Python process that needs to be managed differently. If Python provides a way to inspect memory growth, object counts, thread activity, open resources, or other runtime information that could help compare the application’s state before and after the slowdown develops, I would appreciate recommendations on what to monitor. I would also like to know whether there are specific signs that would distinguish a Python application-level issue from a delay caused by the database, network, or another dependency, since the endpoint itself performs several operations before returning its response.

Has anyone experienced a Python web application where a particular endpoint initially responds normally but becomes intermittently slow after the same application process has been running and handling requests for some time? I would appreciate advice on the most effective way to capture enough diagnostic information during both a fast request and a slow request to identify where the additional time is being introduced. I can provide a simplified version of the endpoint, the Python version, web framework being used, server configuration, request timing information, and relevant log output if those details would help narrow down the cause. For now, I want to keep the investigation focused on this single issue: the Python website’s requests become intermittently slower during a long-running application session, while restarting the application temporarily restores normal response times. I would like to understand how to trace the slowdown to its actual source and determine whether it is related to Python runtime state, application resource management, request concurrency, or another part of the request-processing path.

First guess is a memory leak in your python code.

Second guess is that you have a algorithm that depends on info that you keep from earlier requests and searching that cache is slowing you down.

Which OS are you running this on?