Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

"It's been raining all day and you find yourself with a sufficiently large pile of items (tweets, blog posts, cat pictures) and a key-value database. New items are arriving every minute and you'd really like a way of finding similar items that already exist in your dataset (either for duplication detection or finding related items). Clearly we don't want to scan our entire existing database of items every time we receive a new item but how do we avoid doing so? Minhashing to the rescue!"


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: