# Python Regex and re.findall problems

**URL:** <https://discuss.python.org/t/python-regex-and-re-findall-problems/84310>\
**Category:** Python Help\
**Created:** [March 13, 2025, 11:04am UTC](https://discuss.python.org/t/python-regex-and-re-findall-problems/84310 "2025-03-13T11:04:44Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![c-rob](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/c-rob/32/17261_2.png) [@c-rob](https://discuss.python.org/u/c-rob)\
**Post date:** [March 13, 2025, 11:04am UTC](https://discuss.python.org/t/python-regex-and-re-findall-problems/84310/1 "2025-03-13T11:04:44Z")

</div>

In a string I want to find what we call “macros”, but only some of them, and I want to find all macros that match: \<rx;ABC123\>, or \<grf;ABC144\> or \<grfa;BDB199\>. Here’s the code I’m using and the results. I’ve never done this regex in Python before and it’s not giving me what I want.

```python
'''Test program to test finding multiple macros.'''

import re

lin = '<ps;2><rx;spec><px;;1>Table of contents<pa><spd;1><grf;1212><qa>'
rxarr = re.findall(r'<(grf|grfa|rx);.+?>', lin)
print(rxarr)
# I'm getting ['rx', 'grf'] which is incorrect. 
# I want to get in rxarr: ['<rx;spec>', '<grf;1212>']
rxarr = re.findall(r'(<(grf|grfa|rx);.+?>)', lin)
print(rxarr)
# For this pattern I get rxarr of: [('<rx;spec>', 'rx'), ('<grf;1212>', 'grf')] which I don't want.

```

I wasn’t sure what search terms to use. So how can I use regex to get what I want by using one `re.search()` statement?

Thank you!

1. EDIT: I’m not searching HTML but the strings still look like HTML tags.

---

<div class="post-metadata">

**Author:** ![JamesParrott](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/jamesparrott/32/10534_2.png) [@JamesParrott](https://discuss.python.org/u/JamesParrott)\
**Post date:** [March 13, 2025, 11:25am UTC](https://discuss.python.org/t/python-regex-and-re-findall-problems/84310/2 "2025-03-13T11:25:45Z")

</div>

> **[re — Regular expression operations](https://docs.python.org/3/library/re.html#re.findall)**
>
> Source code: Lib/re/ This module provides regular expression matching operations similar to those found in Perl. Both patterns and strings to be searched can be Unicode strings ( str) as well as 8-...

> The result depends on the number of capturing groups in the pattern. … If there is exactly one group, return a list of strings matching that group.

Try finditer instead and use the entire Match objects to get exacly what you want.

Also what’s the intent behind `.+?` ? If the stuff after the semicolon really is optional, I’d prefer `.*` instead.

---

<div class="post-metadata">

**Author:** ![c-rob](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/c-rob/32/17261_2.png) [@c-rob](https://discuss.python.org/u/c-rob)\
**Post date:** [March 13, 2025, 11:41am UTC](https://discuss.python.org/t/python-regex-and-re-findall-problems/84310/3 "2025-03-13T11:41:50Z")

</div>

> [@c-rob](#):
>
> `r'<(grf|grfa|rx);.+?>'`

This `r'<(grf|grfa|rx);.+?>'` is what limits me to finding several whole macros, not all macros. `.+?>` stops at the first `>` sign. The question mark is a non-greedy modifier to `.+`.

Let’s say I want to find all `<p>` and `<a>` elements in html in this string:  
`<p style="color:red;"><b>My bold text</b> <i>My italic text</i> <a href="https://google.com>Google</a>` We would want to return a list that has only: `['<p style="color:red;">', '<a href="https://google.com>']`

Try this code to see what I mean:

```python
import re
rxarr = re.findall(r'(<(grf|grfa|rx);.*)', lin)
print(rxarr)

```

You can do some testing on [https://regex101.com](https://regex101.com). I just am new to Python regex so I did not remember re.finditer().  
It was probably in a tutorial a year ago but I have never used it before.

Thanks!

---

<div class="post-metadata">

**Author:** ![JamesParrott](https://sea2.discourse-cdn.com/flex002/user_avatar/discuss.python.org/jamesparrott/32/10534_2.png) [@JamesParrott](https://discuss.python.org/u/JamesParrott)\
**Post date:** [March 13, 2025, 11:53am UTC](https://discuss.python.org/t/python-regex-and-re-findall-problems/84310/4 "2025-03-13T11:53:38Z")

</div>

Ah of course, I’d forgot about non-greedy modifiers - thanks
