| Issue |
ITM Web Conf.
Volume 86, 2026
5th International Conference on Current Research in Engineering and Technology (ICCRET-2026)
|
|
|---|---|---|
| Article Number | 01008 | |
| Number of page(s) | 12 | |
| Section | AI & Intelligent Computing | |
| DOI | https://doi.org/10.1051/itmconf/20268601008 | |
| Published online | 05 June 2026 | |
Resume2Role: Intelligent Resume Parsing and Domain Prediction Using NLP and Classification Models
1 Department of Computer Science & Engineering, Amity University, Lucknow
2 Department of Computer Science & Engineering, Amity University, Lucknow
3 Department of Computer Science & Engineering, Amity University, Lucknow
4 Department of Computer Science & Engineering, Brainware University, Kolkata
* e-mail: This email address is being protected from spambots. You need JavaScript enabled to view it.
** e-mail: This email address is being protected from spambots. You need JavaScript enabled to view it.
*** e-mail: This email address is being protected from spambots. You need JavaScript enabled to view it.
**** e-mail: This email address is being protected from spambots. You need JavaScript enabled to view it.
Abstract
Accurate and effective processing of a high volume of resumes poses significant challenges in the dynamic field of digital recruitment. Traditional manual screening methods are highly inefficient, inaccurate, and susceptible to biases. To overcome these limitations, this paper introduces the Resume2Role solution, which utilizes machine learning algorithms and NLP techniques to automatically parse and classify resumes based on job roles. The process involves the use of Python for text extraction from resumes formatted in either TXT or PDF form. The data is cleaned and normalised using preprocessing methods such as lemmatisation, lowercasing, and stop-word removal. Term Frequency-Inverse Document Frequency (TFIDF), which successfully transforms unstructured resume data into structured numerical representations, is then used to vectorise the processed text. Several models were evaluated for classification, and the Random Forest Classifier produced the best accuracy in identifying the appropriate job domain (e.g., IT, HR, Finance, Sales). Further-more, methods such as oversampling were used to address class imbalance, which improved the model’s capacity for generalisation.The creation of a completely functional Flask-based web application, which enables real-time resume uploads, parses important information (such as name, email, phone number, and skills), and provides rapid domain classification with job recommendations, is a significant contribution of this effort. The application performed well in a variety of professional areas and résumé formats. The system’s efficacy is validated by evaluation measures such as confusion matrix analysis, recall, accuracy, and precision.
Key words: Resume Classification / Job Role Prediction / Natural Language Processing / TF-IDF / Random Forest / Web Application / Flask / Machine Learning / Human Resource Automation / Text Mining
© The Authors, published by EDP Sciences, 2026
This is an Open Access article distributed under the terms of the Creative Commons Attribution License 4.0, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Current usage metrics show cumulative count of Article Views (full-text article views including HTML views, PDF and ePub downloads, according to the available data) and Abstracts Views on Vision4Press platform.
Data correspond to usage on the plateform after 2015. The current usage metrics is available 48-96 hours after online publication and is updated daily on week days.
Initial download of the metrics may take a while.

