Skip to content
SmallBizDesk
All tools

Remove duplicate rows from Excel and CSV files

Pick the columns that should be unique, such as Email, and see every repeat before anything is deleted. Keep the first copy or the last.

Remove duplicates

Not uploaded.

Drop the list to clean up

One Excel (.xlsx, .xls), OpenDocument or CSV file. You'll see every duplicate before anything is removed.

How to remove duplicate rows

  1. Open your file. If it's a workbook with several sheets, choose the sheet to clean.
  2. Pick the columns that decide whether two rows are the same. All columns catches rows that were pasted twice; a single column, such as Email, leaves one row per email address.
  3. Check the sets of duplicates on the right. Each one shows the row that stays and the rows that go, with the row numbers you'll see in Excel.
  4. Download the cleaned file. You can also download the removed rows, or keep every row and mark the repeats instead.

Which columns should decide a duplicate?

It depends on what one row stands for. Compare the columns that identify that thing, and leave out the ones that can change between copies, such as a phone number someone updated.

  • Customer or mailing lists: Email. If some people have no email, use first name, last name and ZIP code instead.
  • Orders, invoices or payments: the order or invoice number.
  • An export pasted in twice: all columns.

Rows with nothing in the compared columns are never treated as duplicates of each other. If you compare by Email, ten customers without an email address stay ten customers.

A worked example

A short customer list where two people appear twice, typed slightly differently:

customers.csv

RowABC
1NameEmailCity
2Dana Whitfielddana@example.comColumbus
3Marcus Bellmarcus@example.comAustin
4, removedmarcus bellMARCUS@example.com Austin
5Priya Ramanpriya@example.comTampa
6, removedDana Whitfielddana@example.comDayton
Compared by Email only. Rows 4 and 6 repeat earlier rows once capital letters and spaces are ignored.

Row 6 is a good reminder to look before you download: Dana's city changed. If the newer address is the right one, choose to keep the last copy instead of the first.

Keep the first copy or the last?

Duplicates aren't always identical. When the compared columns match but others differ, the copy you keep decides which version survives.

  • Keep the first when the oldest record is the one that counts, for example the date a customer first signed up.
  • Keep the last when rows are added over time and the newest one has the current details, such as a new address or phone number.

If neither copy is fully right, mark the duplicates instead, fix the kept row by hand, then delete the marked ones.

Reviewing duplicates before deleting them

Mark them only keeps every row and adds a Duplicate of row column. The rows that stay have it empty; each repeat shows the row number it matches.

In Excel, turn on Data > Filter and filter that column to hide the blanks: you see only the repeats, next to their row numbers, and can check them one by one. Sorting by the compared column, such as Email, puts each repeat right under the row it matches. When you're happy, delete the filtered rows and the column.

The same review works for a colleague: send the marked file, and nothing is lost until someone decides.

What doesn't count as a duplicate

Matching is exact once capital letters and extra spaces are set aside. That means some rows a person would call duplicates are not caught:

  • typos, such as dana@exmaple.com
  • nicknames, such as Bob and Robert
  • the same phone number written two ways, such as (614) 555-0142 and 614.555.0142

Guessing at these would also remove rows that only look alike, which is worse than leaving a few in. Comparing by a column that is typed consistently, usually email, avoids most of them.

Removing duplicates in Excel or Google Sheets

Excel has a built-in command: select the table, then Data > Remove Duplicates and tick the columns to compare. It deletes the repeats straight away and tells you how many went, but not which ones, so save a copy first. A trailing space makes two values different to Excel, which is a common reason it seems to miss duplicates; the TRIM function removes those spaces.

To see duplicates without deleting anything, use Home > Conditional Formatting > Highlight Cells Rules > Duplicate Values. It highlights repeated values in the cells you select, one column at a time, rather than whole rows.

In Google Sheets, select the data and use Data > Data cleanup > Remove duplicates.

If your duplicates come from combining several files, merge them first, then clean the merged file here.

Questions

Can I remove duplicates based on one column?

Yes. Pick only that column, for example Email. A whole row is removed when its email repeats an earlier row, and one row is kept for each email address.

Which copy is kept?

The first one in the file, unless you choose to keep the last. Keeping the last is useful when newer rows are added at the bottom.

Can I see what was removed?

Yes. Sets of duplicates are shown before you download, with their row numbers. You can also download the removed rows as a separate file, with a column saying which row each one repeats.

Can I find duplicates without deleting them?

Yes. Choose Mark them only. The file keeps every row and gets a Duplicate of row column showing which row each repeat matches, so you can filter or review them yourself.

Is it case sensitive?

Not by default: Ana@Example.com and ana@example.com count as the same. Untick Ignore capital letters to treat them as different.

Are my files uploaded anywhere?

No. Your browser reads the file on your computer and creates the cleaned copy there too. Nothing is sent to a server.

Last updated by Muhammad Yahya.