Efficient thresholding filter of an array with numpy

Question

I need to filter an array to remove the elements that are lower than a certain threshold. My current code is like this:

threshold = 5
a = numpy.array(range(10)) # testing data
b = numpy.array(filter(lambda x: x >= threshold, a))

The problem is that this creates a temporary list, using a filter with a lambda function (slow).

As this is quite a simple operation, maybe there is a numpy function that does it in an efficient way, but I've been unable to find it.

I thought that another way to achieve this could be sorting the array, finding the index of the threshold and returning a slice from that index onwards, but even if this would be faster for small inputs (and it won't be noticeable anyway), it's definitively asymptotically less efficient as the input size grows.

Update: I took some measurements too, and the sorting + slicing was still twice as fast as the pure python filter when the input was 100.000.000 entries.

r = numpy.random.uniform(0, 1, 100000000)

%timeit test1(r) # filter
# 1 loops, best of 3: 21.3 s per loop

%timeit test2(r) # sort and slice
# 1 loops, best of 3: 11.1 s per loop

%timeit test3(r) # boolean indexing
# 1 loops, best of 3: 1.26 s per loop

yosukesabai · Accepted Answer · 2011-11-03 12:36:44Z

114

b = a[a>threshold] this should do

I tested as follows:

import numpy as np, datetime
# array of zeros and ones interleaved
lrg = np.arange(2).reshape((2,-1)).repeat(1000000,-1).flatten()

t0 = datetime.datetime.now()
flt = lrg[lrg==0]
print datetime.datetime.now() - t0

t0 = datetime.datetime.now()
flt = np.array(filter(lambda x:x==0, lrg))
print datetime.datetime.now() - t0

I got

$ python test.py
0:00:00.028000
0:00:02.461000

http://docs.scipy.org/doc/numpy/user/basics.indexing.html#boolean-or-mask-index-arrays

edited Nov 3, 2011 at 12:36

answered Nov 3, 2011 at 11:55

yosukesabai

6,2144 gold badges32 silver badges42 bronze badges

3

This kind of indexing does not maintain the size of the array, how is it possible to keep the same number of elements and zeroing the subthreshold values?
– linello
Jul 24, 2013 at 10:00
9

@linello, a[a<=threshold] = 0 is going to mask out the part that do not exceed the threshold
– yosukesabai
Aug 17, 2013 at 19:25
4

I ran in to the issue of filtering based on two criteria. Here is the solution: stackoverflow.com/a/3248599/1373468
– Robin Newhouse
Jan 12, 2014 at 3:29
@yosukesabai Is it possible to do exactly this, without actually changing the original values. If np.ma is meant to do that, I cannot figure out how.
– embert
Jan 24, 2014 at 11:04
@embert, not sure what you mean by "changing original values". The array lrg is not changed, and flt has all of values that i wanted (anything but zero)
– yosukesabai
Jan 25, 2014 at 6:11

| Show 6 more comments

cottontail · Accepted Answer · 2023-02-25 00:54:35Z

You can also use np.where to get the indices where the condition is True and use advanced indexing.

import numpy as np
b = a[np.where(a >= threshold)]

One useful function of np.where is that you can use it to replace values (e.g. replace values where the threshold is not met). While a[a <= 5] = 0 modifies a, np.where returns a new array with the same shape only with some values (potentially) changed.

a = np.array([3, 7, 2, 6, 1])
b = np.where(a >= 5, a, 0)       # array([0, 7, 0, 6, 0])

It's also very competitive in terms of performance.

a, threshold = np.random.uniform(0,1,100000000), 0.5

%timeit a[a >= threshold]
# 1.22 s ± 92.2 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)

%timeit a[np.where(a >= threshold)]
# 1.34 s ± 258 ms per loop (mean ± std. dev. of 7 runs, 10 loops each)

Collectives™ on Stack Overflow

Efficient thresholding filter of an array with numpy

2 Answers 2

Not the answer you're looking for? Browse other questions tagged
python
arrays
numpy
filtering
threshold
or ask your own question.

Linked

Hot Network Questions

Collectives™ on Stack Overflow

2 Answers 2

Not the answer you're looking for? Browse other questions tagged pythonarraysnumpyfilteringthreshold or ask your own question.

Linked

Related

Not the answer you're looking for? Browse other questions tagged
python
arrays
numpy
filtering
threshold
or ask your own question.