Analysing the Paris Attacks OSINT data – using SIREN

Paris Attacks

Paris attacks kill at least 128. That was the first news headline I read on my phone on the morning of November 14th 2015. We had our flat in London at the time, it was a sunny day and we were just heading out for brunch at a local café.

BBC News – November 14th 2015

I remember that day well, mainly because we had a friend who was visiting Paris. For the purpose of this blog let’s call this friend ‘Steve’. I quickly checked Steve’s Social Media accounts to see if he was safe; he had no posts since November 10th 2015.

The news reports mentioned the Bataclan and ‘restaurants and bars at five other sites in Paris’. I had no idea where Steve was staying in Paris or indeed his plans for the evening before.

This event occurred when I was working with a commercial Social Media Monitoring platform. This meant I was able to monitor in real-time posts on social media. I decided to monitor keywords such as Paris, Attack, Bomb, Shooting, Shot etc. We then left the flat and I left the monitor running.

While we were at brunch, I checked my personal facebook and saw the following post from Steve:

Although safe, Steve and his friends were confined to their apartment for a while.

Although relieved that Steve was safe 130 people lost their lives during the attack and 413 people were left injured.

After brunch we returned to the flat and I stopped the monitor. I’ve always kept the dataset I generated from that day, it’s a memory of my relief but also of how lucky Steve was.

SIREN – Investigative Intelligence Platform

One of the great things about S-branch, is being software agnostic. We’re not tied to a particular software vendor or piece of software. This means 2 things: the client always gets the best tool for their requirement and we get to play with lots of cool software!

The Siren Investigative Intelligence Platform is something we’ve been playing with for a while now. We’ve been impressed with the demonstrations and tutorials but we really wanted to try it with some real data; like the Paris Attacks data.

We’re not tied to a particular software vendor or piece of software. This means 2 things: the client always gets the best tool for their requirement and we get to play with lots of cool software!

Accessing the data was quick. Siren can analyse data from REST Services, JDBC Data Sources or flat files such as CSV. The ability to use JDBC means that with the correct driver you can connect to pretty much any external database.

The Paris Attacks data was in CSV format. The loading process allowed us to easily perform transformations to the data. We only did some minor formatting transformations, such as splitting a field based on a comma and formatting dates.

SIREN – Loading the data

Autoselect Most Relevant and Generate Dashboard

Once loaded it’s incredibly quick to start gleaming insights from the data. There were two features that we loved: Autoselect Most Relevant and Generate Dashboard.

These two processes took less than a minute to run and automates something which can take a long time to design and get right.

Autoselect Most Relevant analyses the fields from the data source and selects which ones contain the most relevant data for analysis. Generate Dashboard then takes these fields and generates a dashboard from them. These two processes took less than a minute to run and automates something which can take a long time to design and get right.

SIREN – Generate Dashboard

The finished dashboard gives a good idea of what was being mentioned along with where, when and who. This shows just how versatile and quick Siren can be with any sort of data.

Graph Explorer

Dashboards give a great overview of data. It allows analysts to quickly understand and drill down into large sets of data. Once areas of interest have been identified, analyst’s usually want to ‘look into the weeds’ of the data. One way to do this is to look at the rows of data, this can be time consuming and tedious. Another way is to use Graph Visualisations.

Siren has a relations auto-discovery wizard which is in a Beta state at the moment. As data modelling is something S-branch does on a daily basis we chose to do this manually.

Graph Visualisations require data modelling. This is the process where you model entities or nodes and relations or edges. Siren has a relations auto-discovery wizard which is in a Beta state at the moment. As data modelling is something S-branch does on a daily basis we chose to do this manually.

SIREN – Graph Explorer

Once the data had been modelled it was then possible to visualise the results of the social media dashboard on the graph explorer. It is also possible to overlay this information alongside other data from other dashboards. In the example above the Paris Attacks data was narrowed down to posts geotagged within the Paris area. We can clearly see users retweeting the same message.

Summary

This exercise demonstrated to us how versatile Siren can be and how quickly you can glean insights from a set of data. Siren has recently announced the addition of NLP (Natural Language Processing) and Anomaly Detection, meaning we could glean even more information from the data. We look forward to trying this out.

For more information on Siren visit their website or contact S-branch. If you’d like to find out more about visual analysis software, then feel free to call.

S-branch offers independent advice, meaning if Siren is not for you at this time, we can help you find the solution that is.

Visallo – plug and play cognitive technology

VisalloCognitive

Cognitive Technology

In its simplest form, cognitive technology is software mimicking the human brain. A human brain can read text, listen to sound and spot objects in videos. A human brain understands context, it can learn from previous experience and it can adapt as new information comes its way. In short: the human brain is brilliant. Why would we use anything else?

Well, we are in an age of data generation. On average users of the Internet generate 2.5 quintillion bytes of data each day. It’s not just reserved for the Internet either, companies are generating more internal data than ever and as storage becomes cheaper the data only grows.

The key to a company’s success may be hidden in their own data. The meaning of life may well be held on the Internet; you never know.

With this newly acquired data comes great opportunity. The key to a company’s success may be hidden in their own data. The meaning of life may well be held on the Internet; you never know. The problem of course is the sheer volume. The task of understanding the data and obtaining anything useful from it is impossible to do manually; although I’ve seen people try!

This is where cognitive technology comes in. Forms of artificial Intelligence, computer vision, machine learning, natural language processing and speech recognition; that can be used to understand the data for us. It is only then that we can then gain intelligence from the huge amount of data. All that is needed is an analytics system with cognitive technology…

All that is needed is an analytics system with cognitive technology…

Build your own

One option is to build a system from scratch.  By cherry picking the best technology companies (as well as open source resources) that provide cognitive analytics. You could feed data to and from these technologies into the system. You could pick a high-performance graph database to deal with the data and build a cutting edge browser-based UI to give statistical and visual analytics to the users.

This option is a viable but costly one. It will require a good solutions architect who understands the problem, some high-end developers, a data scientist and a project manager. It’s also a time-consuming proposition. You aren’t going to be getting results tomorrow.

Off-the-shelf

Why build something when somebody else has already developed it? Another option is to buy an ‘off the shelf’ solution. Here the deployment will be fast, you will get support when you get stuck and the UI and back end database have already been thought through.

This option is appealing and may be cheaper. It is common though that you are then tied in to any technology that the software vendor has partnered with. This may mean: if you want improved voice recognition or natural language processing in Arabic, you may not find the solution you want.

Visallo – Plug ang Play Cognitive Technology

Visallo is one of the ‘off the shelf’ solutions described above. It is based on a high-performance graph database and it allows users to perform: statistical, visual (graph), temporal and geo-spatial analysis all from one intuitive browser-based UI.

Visallo

Visallo

While that is all appealing, the interesting part to Visallo is their attitude to extension and customization. They actively encourage it. Jeff Kunkle, the President of Visallo actually wrote a blog entitled ‘We love when customers replace our software’

Visallo was designed with extension and customization in mind. From the underlying data store to data processing algorithms and UI plugins, you’re not stuck with a one-size-fits-all solution.

This means if the default OpenNLP plug-in that analyses text and extracts entities (People, Organisations, Locations and Vehicles) from unstructured data doesn’t fit your needs; choose another one. If you want to integrate with you existing investment in voice recognition technology; no problem.

analyticplugin

Plug and Play

This attitude to ‘plug and play’ technology allows the fast deployment and support of an off-the-shelf product, with the flexibility of a build-your-own solution.

For more information on Visallo visit their website or contact S-branch to find out more. If you’d like to find out more about cognitive technologies and the application into crime and fraud; feel free to call.

S-branch offers independent advice, so even if Visallo is not for you, we can help you find the solution that is.