ChatCrawlersearch across public Telegram Open the app
P

Python

snapshot for July 2025
July 2025 ×
31 July 2025
I'm currently using hash duplication detection, but I think it's too slow. I've heard that deduplication can be implemented within the database. I'd like to know which database can do this.
1
Does anyone know how to detect duplication in large amounts of data?
Jürgen Rominsany
I am implementing a quick detection of similarity, which is currently equivalent to checking for plagiarism in papers, but I need to have fewer words. Do you mean that all databases support this? I would like to ask for your advice.
1
I just tried to calculate the hash of the text that needs to be checked for duplicates in the database and then pull it into memory, and then use my python program to check it, but I found that this was too slow.
1
Jürgen Rominsif need to have uniq yes all db supporting it
No, brother, my data is currently in MySQL. In MySQL, I have a field called is_repeat. When a duplicate is detected, this field needs to be updated. What I am thinking now is that as long as the text similarity is greater than 50%, the two will be considered duplicates. Brother, do you have any good ideas?
V
Give me ideas for some interesting projects everyone
I
Hey guys I just started learning python. I need help on this to get it learn
For libraries, statements and recommended youtube videos easy to understand
I
I searched but got confused lots of channels and videos related to python
Open in Telegram Каталог площадок Искать в ChatCrawler

A snapshot of an open public feed from the search index ChatCrawler — “Google for public Telegram”; refreshed as the venue is crawled. Times are UTC.

Public content only, official Telegram API. About · FAQ · What we do not do · Remove a page · Catalog