ChatCrawlersearch across public Telegram Open the app
P

Python

сообщение · 2025-07-31 09:22 UTC
1
Hello, I would like to ask a question. I currently have nearly one million text data, but there is a lot of duplicate data in it. I wonder which database built-in algorithm supports duplicate detection. For example, if the similarity reaches 0.5, it is regarded as two data duplications. Currently, the database is in MySQL. I want to achieve rapid weight-judgment. What should I do?

Вся лента · оригинал в Telegram

Open in Telegram Каталог площадок Искать в ChatCrawler

A snapshot of an open public feed from the search index ChatCrawler — “Google for public Telegram”; refreshed as the venue is crawled. Times are UTC.

Public content only, official Telegram API. About · FAQ · What we do not do · Remove a page · Catalog