PROJECT INDEX
AI

Illegal Content Website Collection and Analysis

InfraBigDataDistributed System
SHEET
009
PERIOD
2023
REV
2026-09-03
READ
1 MIN
LANG
EN
SYSTEM CARD6/7 FIELDS · DRAFTED FROM THE TEXT
01 RUNTIMECelery-based distributed system with containerized architecture
02 DATAover 10,000 daily cases of videos, torrents, webtoons, copyrighted materials from illegal websites
03 MODELcustom AI-powered classification system, text/video frame/metadata identification algorithms, Celery, Selenium, RabbitMQ
04 INTERFACEAPIs, proposed UI features for operators
05 DECISIONtracking and analysis of illegal websites, precise tracking of copyrighted materials
06 FAILUREunknown
07 COST · ENERGYreduced code size by 80%
unknown ≠ n/a

Project Overview

A large-scale project in collaboration with several public and private entities to track and analyze illegal websites.

Screenshots are omitted under the project's license terms.

Main Features

  1. Automated Detection and Collection of Illegal Content
    • Developed a system capable of automatically detecting and collecting over 10,000 cases daily.
    • Covered a wide range of digital content, including videos, torrents, webtoons, and various copyrighted materials.
  2. Significantly Reduced Manual Workload for Operators
    • Implemented a custom AI-powered classification system to assist operators in analyzing illegal content.
    • Automated decision-making processes reduced the need for human intervention, improving efficiency and accuracy.
  3. AI-Based Content Identification for Copyright Protection
    • Developed text-based, video frame-based, and metadata-based identification algorithms.
    • Enabled precise tracking of specific copyrighted materials, even across modified or re-encoded formats.
  4. Fully Automated Workflow with Celery, Containers, and Selenium
    • Designed a Celery-based distributed system for scalable data collection and processing.
    • Leveraged containerized architecture for flexible deployment and seamless scalability.
    • Automated data scraping, classification, and reporting using Selenium and other web automation tools.

Key Contributions

  • Conducted preliminary research on AI-based web scraping technologies.
  • Drafted proposal documents and presentation materials after analyzing requirements.
  • Developed a distributed data collection system using RabbitMQ and Python Celery.
  • Created algorithms for automatically tracking hundreds of illegal website domains.
  • Designed APIs to check the visibility of illegal websites on major search engines.
  • Proposed new UI features and designed database schemas (ERD).
  • Implemented PostgreSQL integration using Peewee ORM.
  • Refactored legacy service code, reducing code size by 80%.

Achievements

  • Contributed to winning project bids and took a lead role in ensuring code quality for one of the two teams.
KADE-6 · ANSWERS FROM THIS AND OTHER SHEETSSOURCES: 16 POSTS

› Ask about this sheet. Answers cite sheet numbers.