Skip to content
Back to all projects
  • Machine Learning
  • NLP

Code-Mixed Humorous Text Detection

An NLP system that detects humor in code-mixed Telugu-English text.

Overview

A custom dataset of 5,000 romanized Telugu-English code-mixed jokes is curated and used to train a humor classification model, augmented with supplementary English joke datasets and evaluated on accuracy, precision, recall and F1-score.

Problem Statement

Almost all humor detection research assumes clean, monolingual English. Indian social media is overwhelmingly code-mixed and romanized, which breaks both tokenizers and pretrained embeddings.

Proposed Solution

Build the dataset that does not exist, then train and evaluate humor classification directly on romanized Telugu-English rather than translating it into English first.

Key Features

  • Custom Dataset Curation

    5,000 romanized Telugu-English code-mixed jokes.

  • Humor Classification Model

    Detects humorous vs non-humorous text.

  • Cross-Lingual Dataset Augmentation

    Supplementary English joke datasets for training.

  • Performance Evaluation

    Accuracy, precision, recall, and F1-score reporting.

Technologies Used

NLP

  • Code-mixed tokenisation
  • TF-IDF
  • Word embeddings
  • NLTK

Machine Learning

  • Scikit-Learn
  • SVM
  • Logistic Regression

Data

  • Custom corpus curation
  • Pandas
  • Annotation pipeline

Project Workflow

  1. 1

    Dataset curation

    Romanized Telugu-English jokes are collected and labelled.

  2. 2

    Preprocessing

    Code-mixed text is normalised and tokenised.

  3. 3

    Augmentation

    English joke corpora supplement the training data.

  4. 4

    Classification

    The model predicts humorous vs non-humorous.

  5. 5

    Evaluation

    Results are reported across standard metrics.

What You'll Receive

  • Complete, runnable source code with folder structure
  • Project report and technical documentation
  • Ready-to-present PPT content
  • Setup and installation walkthrough
  • Line-by-line project explanation session
  • Viva question bank with answers
  • Bug fixing and troubleshooting help
  • Post-delivery support after submission

Frequently Asked Questions

How much does this project cost?

Every project is quoted individually, because the price depends on the modules you need, your technology stack, your college's format and your deadline. Send us the project name on WhatsApp and we'll share a quote the same day.

Can the project be customised to my college requirements?

Yes. Share your guide's requirements, preferred technology stack and abstract format, and we adapt the modules, dataset or UI accordingly.

Will I be able to explain this project during my viva?

That is the point of the explanation session. We walk you through the architecture, every module, the flow of data and the results, and hand over a viva question bank with answers.

What if the project does not run on my laptop?

We help you with setup end to end — dependencies, environment, database and configuration — over chat or a call until it runs on your machine.

Do I get support after submission?

Yes. Post-delivery support is included, so you can come back for fixes, doubts or demo help even after the project is delivered.

Want This Project?

Get complete source code, documentation, PPT, explanation, and support.

Similar projects

  • Blockchain
  • Gen-AI
  • NLP

Blockchain-Based Complaint Management System

A secure, intelligent complaint management platform ensuring transparency, integrity, and automation.

  • IPFS-Based Storage
  • Blockchain-Based Transaction Log
  • GenAI Text Analysis
  • +1 more

Custom quote

Priced to your scope

View Project
  • AI
  • NLP

AI-Powered Interview Simulator

An end-to-end interview preparation platform leveraging AI for realistic practice and evaluation.

  • Resume-Based Screening Engine
  • Voice-Based Interview Practice
  • Auto-Generated Assessments

Custom quote

Priced to your scope

View Project