# Add glob.translate(): convert path with shell wildcards to regular expression

**URL:** <https://discuss.python.org/t/add-glob-translate-convert-path-with-shell-wildcards-to-regular-expression/31549>\
**Category:** Ideas\
**Created:** [August 13, 2023, 11:33am UTC](https://discuss.python.org/t/add-glob-translate-convert-path-with-shell-wildcards-to-regular-expression/31549 "2023-08-13T11:33:09Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

**Author:** ![barneygale](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/barneygale/32/1882_2.png) [@barneygale](https://discuss.python.org/u/barneygale)\
**Post date:** [August 13, 2023, 11:33am UTC](https://discuss.python.org/t/add-glob-translate-convert-path-with-shell-wildcards-to-regular-expression/31549/1 "2023-08-13T11:33:09Z")

</div>

Quoth the [`fnmatch`](https://docs.python.org/3/library/fnmatch.html) docs:

> Note that the filename separator (`'/'` on Unix) is _not_ special to this module. See module [`glob`](https://docs.python.org/3/library/glob.html#module-glob) for pathname expansion ([`glob`](https://docs.python.org/3/library/glob.html#module-glob) uses [`filter()`](https://docs.python.org/3/library/fnmatch.html#fnmatch.filter) to match pathname segments).

So `fnmatch` operates only on _filenames_. What if we want to translate, match or filter on _paths_? The table below shows the situation:

| | arg: **filename** | arg: **path** |
| --- | --- | --- |
| **translate** | `fnmatch.translate()` | _not supported!_ |
| **match** | `fnmatch.fnmatch()` `fnmatch.fnmatchcase()` | `pathlib.PurePath.match()` |
| **filter** | `fnmatch.filter()` | _not supported!_ |
| **glob** | _N/A_ | `pathlib.Path.glob()` `glob.glob()` |

[GH-72904](https://github.com/python/cpython/issues/72904) requests that we fill in that top right “not supported” cell.

I propose we add a `glob.translate()` function that converts a _path_ with shell-style wildcards to a regular expression. I have an implementation available in [GH-106703](https://github.com/python/cpython/pull/106703), which also adjusts pathlib to call the new function for a tidy speedup.

Thoughts? Qs from my side:

1. Does this seem sensible?
2. Should the function support a _recursive_ argument, like `glob()`? Should we match its default (false)?
3. Should the function support an _include\_hidden_ argument, like `glob()`? Should we match its default (false)?
4. How worried should I be about [exponential execution time](https://github.com/python/cpython/issues/84660)? IIUC [the fix for this in `fnmatch.translate()`](https://github.com/python/cpython/pull/19908) won’t carry over to `glob.translate()`, because it relies on all the variable-width parts matching anything. (I may be totally wrong here.)
5. Am I alright to copy-paste parts of the `fnmatch.translate()` implementation, particularly `[seq]` handling, which is common to both versions? Or should I look adding some sort of `fnmatch.Translator` class that can be subclassed in `glob`? Or something else?

Ta!

---

<div class="post-metadata">

**Author:** ![facelessuser](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/facelessuser/32/1971_2.png) [@facelessuser](https://discuss.python.org/u/facelessuser)\
**Post date:** [August 13, 2023, 7:45pm UTC](https://discuss.python.org/t/add-glob-translate-convert-path-with-shell-wildcards-to-regular-expression/31549/2 "2023-08-13T19:45:31Z")

</div>

I personally think a `glob.translate()` seems sensible, but then again, I’ve written a dedicated library for this sort of thing, `glob.translate()` being [here](https://facelessuser.github.io/wcmatch/glob/#translate). I think it can be useful if you want to generate some hard path matches for a script, but don’t want it to have any dependencies.

I think being able to include and exclude hidden files is useful as well.

I probably can’t recommend how it should be implemented in the existing Python framework though. Happy to see some of these ideas making it into the standard lib though.

---

<div class="post-metadata">

**Author:** ![facelessuser](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/facelessuser/32/1971_2.png) [@facelessuser](https://discuss.python.org/u/facelessuser)\
**Post date:** [August 13, 2023, 8:15pm UTC](https://discuss.python.org/t/add-glob-translate-convert-path-with-shell-wildcards-to-regular-expression/31549/3 "2023-08-13T20:15:05Z")

</div>

I guess I was thinking about this from an external library perspective. I guess if this is in the standard lib, it can still be useful as you can generate the pattern once, compile it, and not waste time regenerating the pattern and compiling it in the future.

---

<div class="post-metadata">

**Author:** ![barneygale](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/barneygale/32/1882_2.png) [@barneygale](https://discuss.python.org/u/barneygale)\
**Post date:** [September 30, 2023, 8:55pm UTC](https://discuss.python.org/t/add-glob-translate-convert-path-with-shell-wildcards-to-regular-expression/31549/4 "2023-09-30T20:55:25Z")

</div>

Update: I’ve added _recursive_ and _include\_hidden_ arguments to my implementation in [GH-106703](https://github.com/python/cpython/pull/106703), so it should match `glob.glob()` exactly. @jaraco has approved an earlier version, but it would be good to get some eyes on the latest version. Would anyone be up for reviewing? Thanks!

---

<div class="post-metadata">

**Author:** ![barneygale](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/barneygale/32/1882_2.png) [@barneygale](https://discuss.python.org/u/barneygale)\
**Post date:** [December 1, 2023, 10:43pm UTC](https://discuss.python.org/t/add-glob-translate-convert-path-with-shell-wildcards-to-regular-expression/31549/5 "2023-12-01T22:43:29Z")

</div>

For posterity: [`glob.translate()`](https://docs.python.org/3.13/library/glob.html#glob.translate) has landed in 3.13. Thanks for your pointers @facelessuser, and thanks also folks who helped review [the PR](https://github.com/python/cpython/pull/106703).
