Python - Python Sets and Set Operations in Depth
Introduction
A set in Python is a built-in collection used to store multiple values where every element must be unique. Unlike lists and tuples, sets do not maintain duplicate values. Sets are particularly useful when you need to perform mathematical set operations such as union, intersection, difference, and symmetric difference.
Sets are mutable, which means that elements can be added or removed after the set has been created. However, the elements stored inside a set must be hashable, meaning immutable types such as integers, strings, and tuples can generally be stored, while mutable objects such as lists and dictionaries cannot be directly stored in a set.
A set is created by placing elements inside curly braces {} or by using the set() constructor.
numbers = {10, 20, 30, 40}
print(numbers)
Output:
{40, 10, 20, 30}
The order of elements should not be relied upon because sets are unordered collections.
Creating a Set
A set can be created by using curly braces:
fruits = {"apple", "banana", "orange"}
print(fruits)
A set can also be created from another iterable using set():
numbers = set([10, 20, 30, 40])
print(numbers)
The set() constructor is especially useful when converting a list, tuple, or another iterable into a set.
values = [10, 20, 20, 30, 30, 40]
unique_values = set(values)
print(unique_values)
Output:
{10, 20, 30, 40}
Here, duplicate values are automatically removed.
Creating an Empty Set
An important point is that {} does not create an empty set. It creates an empty dictionary.
data = {}
print(type(data))
Output:
<class 'dict'>
To create an empty set, use set():
data = set()
print(type(data))
Output:
<class 'set'>
This distinction is important when working with Python collections.
Duplicate Elements in a Set
Sets automatically eliminate duplicate values.
numbers = {1, 2, 2, 3, 3, 4}
print(numbers)
Output:
{1, 2, 3, 4}
This makes sets useful when you need to remove duplicate data.
For example:
names = ["John", "Mary", "John", "David", "Mary"]
unique_names = set(names)
print(unique_names)
The resulting set contains each name only once.
Adding Elements to a Set
The add() method is used to add a single element.
languages = {"Python", "Java"}
languages.add("C++")
print(languages)
The new element is added to the set.
If the element already exists, add() does not create a duplicate.
languages = {"Python", "Java"}
languages.add("Python")
print(languages)
The set still contains only one "Python" element.
Adding Multiple Elements
The update() method can be used to add multiple elements.
numbers = {1, 2, 3}
numbers.update([4, 5, 6])
print(numbers)
It can accept different iterable objects:
numbers.update((7, 8))
numbers.update({9, 10})
After these operations, all the new elements are included in the set.
Removing Elements
Python provides several methods for removing elements from a set.
Using remove()
The remove() method removes a specified element.
numbers = {10, 20, 30}
numbers.remove(20)
print(numbers)
If the specified element does not exist, remove() raises a KeyError.
numbers.remove(50)
This produces an error because 50 is not present.
Using discard()
The discard() method also removes an element, but it does not raise an error if the element is absent.
numbers = {10, 20, 30}
numbers.discard(50)
print(numbers)
The program continues normally.
The main difference is:
remove() -> raises an error if the element is absent
discard() -> does not raise an error if the element is absent
Using pop()
The pop() method removes and returns an arbitrary element from the set.
numbers = {10, 20, 30}
value = numbers.pop()
print(value)
print(numbers)
Because sets are unordered, you should not assume which element will be removed.
Using clear()
The clear() method removes all elements.
numbers = {10, 20, 30}
numbers.clear()
print(numbers)
Output:
set()
Checking Whether an Element Exists
The in operator can be used to check whether an element is present.
languages = {"Python", "Java", "C++"}
if "Python" in languages:
print("Python is available")
The not in operator checks whether an element is absent.
if "Ruby" not in languages:
print("Ruby is not available")
Set membership testing is generally efficient because sets are implemented using hash-based data structures.
Union of Sets
The union of two sets contains all unique elements from both sets.
Suppose:
A = {1, 2, 3}
B = {3, 4, 5}
The union is:
{1, 2, 3, 4, 5}
It can be performed using the | operator:
result = A | B
print(result)
It can also be performed using union():
result = A.union(B)
print(result)
Union is useful when combining unique values from multiple collections.
For example, if two departments have lists of employees:
department_a = {"John", "Mary", "David"}
department_b = {"David", "Sarah", "Robert"}
employees = department_a | department_b
print(employees)
The resulting set contains all employees without duplicate names.
Intersection of Sets
The intersection contains only elements that are common to both sets.
A = {1, 2, 3, 4}
B = {3, 4, 5, 6}
result = A & B
print(result)
Output:
{3, 4}
The intersection() method can also be used:
result = A.intersection(B)
Intersection is useful when finding common data.
For example:
python_students = {"Anil", "Ravi", "Meena", "John"}
java_students = {"John", "Meena", "David"}
common_students = python_students & java_students
print(common_students)
The result identifies students who study both subjects.
Difference Between Sets
The difference operation returns elements that are present in the first set but not in the second set.
A = {1, 2, 3, 4}
B = {3, 4, 5, 6}
result = A - B
print(result)
Output:
{1, 2}
This means 1 and 2 exist in A but not in B.
The reverse operation produces a different result:
result = B - A
print(result)
Output:
{5, 6}
The difference() method can also be used:
result = A.difference(B)
Symmetric Difference
The symmetric difference returns elements that belong to either set but not to both sets.
Consider:
A = {1, 2, 3, 4}
B = {3, 4, 5, 6}
The common elements are 3 and 4. Therefore, the symmetric difference contains:
{1, 2, 5, 6}
Using the ^ operator:
result = A ^ B
print(result)
It can also be written as:
result = A.symmetric_difference(B)
This operation is useful when you need to identify values that are unique to either one of two groups.
Subset
A set is a subset of another set if every element in the first set is also present in the second set.
A = {1, 2}
B = {1, 2, 3, 4}
print(A.issubset(B))
Output:
True
The <= operator can also be used:
print(A <= B)
A proper subset can be checked using <:
print(A < B)
This returns True when A is a subset of B but the two sets are not equal.
Superset
A set is a superset when it contains every element of another set.
A = {1, 2, 3, 4}
B = {1, 2}
print(A.issuperset(B))
Output:
True
The >= operator can also be used:
print(A >= B)
A proper superset can be checked using >.
Disjoint Sets
Two sets are called disjoint when they have no elements in common.
A = {1, 2, 3}
B = {4, 5, 6}
print(A.isdisjoint(B))
Output:
True
If the sets have at least one common element, the result is False.
A = {1, 2, 3}
B = {3, 4, 5}
print(A.isdisjoint(B))
Output:
False
Frozen Sets
Python also provides an immutable version of a set called a frozenset.
numbers = frozenset([1, 2, 3, 4])
print(numbers)
A frozenset cannot be modified after creation. Therefore, methods such as add() and remove() cannot be used on it.
numbers.add(5)
This results in an error.
Frozensets are useful when you need set-like behavior but want the collection to remain unchanged. Because a frozenset is immutable and hashable, it can also be used as an element of another set.
A = frozenset([1, 2, 3])
B = {A}
print(B)
Set Comprehension
Python supports set comprehensions, which provide a concise way to create sets.
For example:
squares = {x * x for x in range(1, 6)}
print(squares)
Output:
{1, 4, 9, 16, 25}
A condition can also be included:
even_numbers = {x for x in range(1, 11) if x % 2 == 0}
print(even_numbers)
Output:
{2, 4, 6, 8, 10}
Set comprehensions are useful when transforming or filtering data while automatically eliminating duplicates.
Set Operations with Multiple Sets
Python allows operations involving more than two sets.
A = {1, 2, 3}
B = {3, 4, 5}
C = {5, 6, 7}
result = A | B | C
print(result)
The union contains all unique elements.
Similarly:
result = A & B & C
returns elements common to all three sets.
Practical Example
Consider three groups of students:
python = {"Anil", "Ravi", "Meena", "John"}
java = {"John", "Meena", "David", "Ravi"}
cpp = {"Ravi", "David", "Kiran"}
To find all students studying at least one subject:
all_students = python | java | cpp
print(all_students)
To find students studying both Python and Java:
python_java = python & java
print(python_java)
To find students studying Python but not Java:
only_python = python - java
print(only_python)
To find students who study Python or Java but not both:
one_of_two = python ^ java
print(one_of_two)
This demonstrates how set operations can simplify tasks involving groups and relationships.
Advantages of Sets
Sets provide several important advantages in Python.
First, they automatically eliminate duplicate values. This makes them convenient for identifying unique items.
Second, membership testing is generally very efficient:
if item in my_set:
...
Third, Python provides built-in mathematical operations such as union, intersection, difference, and symmetric difference.
Fourth, sets are useful for comparing groups of data without requiring lengthy loops.
Limitations of Sets
Sets also have some limitations. They do not support indexing in the same way lists do.
For example:
numbers = {10, 20, 30}
print(numbers[0])
This results in an error because sets do not provide positional indexing.
If ordered access is required, a list may be more appropriate.
Another limitation is that set elements must be hashable. Therefore, the following is invalid:
data = {[1, 2], [3, 4]}
Lists cannot be directly stored as set elements because lists are mutable and unhashable.
Set vs List
A list allows duplicate values and supports indexing:
numbers = [10, 20, 20, 30]
print(numbers[0])
A set removes duplicates and does not provide positional indexing:
numbers = {10, 20, 20, 30}
print(numbers)
The choice depends on the requirement. Use a list when order and positional access are important. Use a set when uniqueness and membership testing are the main requirements.
Conclusion
Python sets are an important collection type for working with unique data and relationships between groups. They automatically remove duplicates and provide efficient operations for checking membership. Their mathematical operations, including union, intersection, difference, and symmetric difference, make them especially useful for comparing and combining collections.
Understanding methods such as add(), update(), remove(), discard(), and clear(), along with concepts such as subsets, supersets, disjoint sets, frozensets, and set comprehensions, provides a strong foundation for using sets effectively in Python programs.