Textual data
Textual data analysis is an approach that treats texts as data and allows you to explore and visualize collections of various texts. It applies to both short texts (such as comments or tweets) and longer texts (such as articles, reports or books).
Today, the SMCS team has developed expertise in textual data and can therefore assist you in exploring R, Python and Orange software, as well as specific methods such as:
- Web scraping: extracting data from websites
- Text mining: extracting useful information and patterns from unstructured textual data using natural language processing, machine learning or statistical techniques
- Wordcloud: visualization of the most frequently used words in a corpus
- Topic modeling: classification of texts into automatically determined thematic groups
- Sentiment analysis: classification of texts according to their tone (positive, neutral, negative), for example on tweets or customer reviews
- Similarity detection: grouping of similar texts, detection of plagiarism or similar opinions
- Natural language processing (NLP)*: use of a set of techniques and methods to enable a machine to understand, interpret and generate human language - NLP combines fields such as linguistics, computer science and artificial intelligence
(*) In the field of natural language processing (NLP), the SMCS benefits from the extensive expertise of
CENTAL, another UCLouvain platform, with which it has enjoyed a fruitful collaboration for many years.