VA Insight: Filling the hole left by Intellishare

hole left by Intellishare

Intellishare was a lightweight browser based application. It provided searching and basic visualisations across data held in an i2 data repository; in this case i2 iBase.

I can remember when Intellishare was first released. It was the brainchild of some very clever people in the i2 consultancy team. ‘IBase Web’ as it was called then was trialled in Hertfordshire Police here in the UK, before being made available as the product known as Intellishare.

Intellishare became incredibly popular with the i2 community, with its easy to use web-based interface and straightforward deployment model. Its find and explore module gave access to valuable data sitting within an iBase database. While its data entry module, meant that real time information could be captured by front line staff.

Software Withdrawal

In September 2019 Intellishare was withdrawn and support was discontinued for Intellishare. While this was understandable, as a lot of the supporting frameworks were now themselves out of date, this left Intellishare users with some limited options.

Introducing VA Insight Portal

The VA Insight Portal ™ is a modern, 3-tier architecture, browser-based solution allowing users to streamline data collection and to view tactical visualisations while on the go. It offers a simple, cost-effective deployment model that compliments an existing i2 environment.

S-branch have had the opportunity to thoroughly test the latest release of the VA Insight Portal and to share our user experience directly with the Insight development team. Our client base has felt the hole left by Intellishare being discontinued. We feel that VA Insight fills that hole nicely!

We deployed the VA Insight Portal™ on a range of our test SQL source databases. In this blog, we will share our user experience with the platform:

Dashboards

The team at S-branch were very excited to see dashboards in VA Insight. It’s too often that product demos of i2 start with a single entity, a target person or phone for example. The real problem is often identifying that target, which then gives a start point to begin an investigation.

The dashboards are basic, don’t expect a full business intelligence suite like PowerBI or Cognos. However, for getting a “lie of the land” it works well.

Entities and Links

Entities and links can be searched, either by a basic or advanced search. The record view shows a basic visualisation and all the record details in one easy to read screen. One feature that we loved was the export functionality. It allowed a pdf of the record to be created with a visualisation attached; a really nice touch.

As of today, there is no expand function (that you would find in iBase). It is possible to drill down into an entity by clicking on it, this will open up that entity and show its visualisation. We think this is fine for the use cases that VA Insight is targeted at. Hardcore analyst’s will be still using the back office i2 suite, while front line staff will be accessing the data through the easier to use VA Insight.

Queries

We were impressed that VA Insight came with its own visual query builder. We were even more impressed that it supports parameterised queries, such as @#NOWDATE. While this is actually an improvement to the functionality within Intellishare, it should be noted that queries can be made diretly within the source SQL  database and then made available to VA Insight.

Data Entry

VA Insight supports data entry through its entities and links, its visualisation screen and through custom forms. We like the flexibility this brings, and we found the data entry to be intuitive.

Summary

In summary we found the VA Insight Portal™ to be an excellent alternative for IBM i2 Intellishare users. While not a direct replacement, it offers intuitive workflows, increased functionality and it is complimentary to most clients existing infrastructure.

There are other options available to address the hole that was left behind by Intellishare. However, we feel that this is the most appropriate in regard to technology used, the deployment time and importantly the price.

If you’re looking for an Intellishare replacement, looking into i2 or you’re looking into data analysis tools in general; please feel free to contact us.

Analysing the Paris Attacks OSINT data – using SIREN

Paris Attacks

Paris attacks kill at least 128. That was the first news headline I read on my phone on the morning of November 14th 2015. We had our flat in London at the time, it was a sunny day and we were just heading out for brunch at a local café.

BBC News – November 14th 2015

I remember that day well, mainly because we had a friend who was visiting Paris. For the purpose of this blog let’s call this friend ‘Steve’. I quickly checked Steve’s Social Media accounts to see if he was safe; he had no posts since November 10th 2015.

The news reports mentioned the Bataclan and ‘restaurants and bars at five other sites in Paris’. I had no idea where Steve was staying in Paris or indeed his plans for the evening before.

This event occurred when I was working with a commercial Social Media Monitoring platform. This meant I was able to monitor in real-time posts on social media. I decided to monitor keywords such as Paris, Attack, Bomb, Shooting, Shot etc. We then left the flat and I left the monitor running.

While we were at brunch, I checked my personal facebook and saw the following post from Steve:

Although safe, Steve and his friends were confined to their apartment for a while.

Although relieved that Steve was safe 130 people lost their lives during the attack and 413 people were left injured.

After brunch we returned to the flat and I stopped the monitor. I’ve always kept the dataset I generated from that day, it’s a memory of my relief but also of how lucky Steve was.

SIREN – Investigative Intelligence Platform

One of the great things about S-branch, is being software agnostic. We’re not tied to a particular software vendor or piece of software. This means 2 things: the client always gets the best tool for their requirement and we get to play with lots of cool software!

The Siren Investigative Intelligence Platform is something we’ve been playing with for a while now. We’ve been impressed with the demonstrations and tutorials but we really wanted to try it with some real data; like the Paris Attacks data.

We’re not tied to a particular software vendor or piece of software. This means 2 things: the client always gets the best tool for their requirement and we get to play with lots of cool software!

Accessing the data was quick. Siren can analyse data from REST Services, JDBC Data Sources or flat files such as CSV. The ability to use JDBC means that with the correct driver you can connect to pretty much any external database.

The Paris Attacks data was in CSV format. The loading process allowed us to easily perform transformations to the data. We only did some minor formatting transformations, such as splitting a field based on a comma and formatting dates.

SIREN – Loading the data

Autoselect Most Relevant and Generate Dashboard

Once loaded it’s incredibly quick to start gleaming insights from the data. There were two features that we loved: Autoselect Most Relevant and Generate Dashboard.

These two processes took less than a minute to run and automates something which can take a long time to design and get right.

Autoselect Most Relevant analyses the fields from the data source and selects which ones contain the most relevant data for analysis. Generate Dashboard then takes these fields and generates a dashboard from them. These two processes took less than a minute to run and automates something which can take a long time to design and get right.

SIREN – Generate Dashboard

The finished dashboard gives a good idea of what was being mentioned along with where, when and who. This shows just how versatile and quick Siren can be with any sort of data.

Graph Explorer

Dashboards give a great overview of data. It allows analysts to quickly understand and drill down into large sets of data. Once areas of interest have been identified, analyst’s usually want to ‘look into the weeds’ of the data. One way to do this is to look at the rows of data, this can be time consuming and tedious. Another way is to use Graph Visualisations.

Siren has a relations auto-discovery wizard which is in a Beta state at the moment. As data modelling is something S-branch does on a daily basis we chose to do this manually.

Graph Visualisations require data modelling. This is the process where you model entities or nodes and relations or edges. Siren has a relations auto-discovery wizard which is in a Beta state at the moment. As data modelling is something S-branch does on a daily basis we chose to do this manually.

SIREN – Graph Explorer

Once the data had been modelled it was then possible to visualise the results of the social media dashboard on the graph explorer. It is also possible to overlay this information alongside other data from other dashboards. In the example above the Paris Attacks data was narrowed down to posts geotagged within the Paris area. We can clearly see users retweeting the same message.

Summary

This exercise demonstrated to us how versatile Siren can be and how quickly you can glean insights from a set of data. Siren has recently announced the addition of NLP (Natural Language Processing) and Anomaly Detection, meaning we could glean even more information from the data. We look forward to trying this out.

For more information on Siren visit their website or contact S-branch. If you’d like to find out more about visual analysis software, then feel free to call.

S-branch offers independent advice, meaning if Siren is not for you at this time, we can help you find the solution that is.

Why Rosoka is the ‘no-brainer’ plug-in for IBM i2 Analyst’s Notebook

Why Rosoka is the ‘no brainer’ plug-in for IBM i2 Analyst’s Notebook

It’s no secret that i2® is a popular choice for visualising and analysing data. That’s why in 2011 IBM paid a rumoured $500 million to acquire the small British software company. Its flagship product: Analyst’s Notebook has now become the standard for charting data and is used in nearly all of the UK police forces. If you have a sporadic flow of data in formats that vary greatly, IBM i2 is still one of the most compelling tools on the market.

IBM i2® products have always worked well with structured data. This has given its users the ability to ‘bring to life’ data hidden in columns and rows. There are many other places that data can hide though and one of those places is in unstructured text.

Unstructured Data

Unstructured text can pose a challenge to analyse and understand, and we keep generating more of it every day. This text can be anywhere from documents, emails or social media posts. In March 2018 we wrote a blog detailing how data hidden in unstructured text can be utilised. It focussed on the use of Natural Language Processing and particularly Entity Extraction. You can read it here. That blog post concentrated on what could be done using freely available tools.

At the time we stated that there are many commercial offerings that offer that functionality but in a much slicker and easy to access interface. One of those tools is Rosoka Text Analytics. Rosoka is an industry leader of text analytics solutions. Their enterprise solution Rosoka Server is great for large scale analysis of unstructured data. This blog is about their standalone solution which is a plug-in for i2.

Rosoka

Rosoka Text Analytics

So, why is it a ‘no-brainer’?

Well, to start off with its fully integrated with IBM i2 Analyst’s Notebook. It’s a plug-in that has been created in close collaboration with IBM and you can tell. We tested the plug-in against several datasets, including the case study we used in the previous blog post. This data can be found here.

Here are the reasons why we think it’s a ‘no-brainer’ plug-in:

  • It extracts entities and links from unstructured files such as PDF’s, Word documents and txt files. It then allows users to generate Analyst’s Notebook charts based on the content
  • Processing time is quick compared to reading the documents manually and quicker than the free extraction engines used in our previous blog
  • During testing we found the entity extraction to be very accurate. What was more impressive was the relationship extraction. In our previous blog we inferred links based on how many times two entities are mentioned in the same document. What Rosoka do is far more sophisticated and is based on the language around the two entities
  • Entity extraction is good but it can create a lot of noise. Rosoka’ s use of Salience (relevance) means we could always find the most relevant extracted entities and ignore the others
  • Rosoka applies entity resolution to its extraction. This meant that if Theresa May is mentioned as ‘Theresa’, ‘Mrs May’ or even ‘She’ in a document, Rosoka will be able to resolve these different texts as one entity.
  • It’s easy to reclassify an extracted entity from say a person to a company (for example Robert Dyas the UK hardware store). Rosoka can also learn from your decisions so it doesn’t make the same mistake again
  • Rosoka is truly multilingual. We tested this by running several news articles about Islamic State through Rosoka in different languages. The organisation Islamic State was mentioned in both the Russian and the Arabic news articles. On the chart they were resolved into one entity, despite them originating from two different languages. That meant that when we expanded Islamic state we got linked items from both the Russian and the Arabic articles.
  • The most surprising reason Rosoka is a no-brainer though is the price. For a fraction of the cost of IBM i2 Analyst’s Notebook the plug-in delivers accurate, multilingual, unstructured data analysis. We’ve looked at costs for this type of technology for clients before and often the requirements were dropped, due to budget restrictions; well not with Rosoka!

In conclusion, we believe Rosoka is a ‘no-brainer’ plug-in for i2 Analyst’s Notebook. The functionality, ease of use and price will just make any analysts life easier!

For more information on Rosoka visit their website or contact S-branch. If you’d like to find out more about unstructured data analysis, then feel free to call.

S-branch offers independent advice, meaning if Rosoka is not for you at this time, we can help you find the solution that is.

Who ya gonna call? A Forensic Accountant…

Forensic Accounting Software

A lot of people won’t know what a forensic accountant is. I didn’t. It was only when I started working with a team of them at a prestigious accountancy firm that I really got to appreciate the importance of the role and the people that do it.

A forensic account is someone who applies accountancy skills to investigate financial discrepancies and inaccuracies. These investigations will be used to find fraudulent activity, financial misrepresentation or even misconduct.

There is one word I would use when describing forensic accountants and that is versatile

Who ya gonna call?

The truth is a forensic accountant could be useful to a lot of us.

If you run a business, for example, you may believe that there is some internal fraud going on – who ya gonna call? A forensic accountant. You may be about to buy a business and want to know more about its books – who ya gonna call? A forensic accountant.

On a personal level you and your siblings may have been left money or maybe you are going through a divorce and believe money is being hidden from you. That’s right, a forensic accountant can help with these things.

Business evaluations, tax audits, insurance term reviews, stolen paper records; there are a lot of situations where a forensic accountant can help.

Skills

There is one word I would use when describing forensic accountants and that is versatile. I’ve seen them successfully deal with sporadic and often unhelpful data to help a client. This is due to the diverse range of investigations they may conduct.

Applying technology to help clients make sense of sporadic and often unhelpful data is what S-branch does best.

Alongside obvious structured data sources such as companies house records, a forensic accountant may receive data in all manner of ways: boxes full of documents, ceased laptops or mobile phones, ancient financial records, interview videos and recordings to name but a few.

How S-branch can help

S-branch

Trusted Analytics Consultancy

Applying technology to help clients make sense of sporadic and often unhelpful data is what S-branch does best. S-branch is truly independent and remains software agnostic, ensuring our clients get the best solution for their needs; whether they need forensic accounting software or just somewhere to store their intelligence.

Here are some of the ways S-branch can help forensic accountants.

Visualisation Software

Finding hidden connections between different data sources, whether that be financial transactions or social media, requires a good eye. This becomes especially important when you are trying to ‘follow the money’.

The use of visualisation analysis tools or graph visualisations to understand complex data is now common among analysts. Otherwise unseen connections between data can be identified and drilled down upon with ease. Visualisation tools also offer a clear way to present findings and share results.

Unstructured Data Analysis

Large amounts of text can pose a challenge to analyse and understand. We keep generating more of it every day. This ‘unstructured text’ can be in documents, emails or social media posts. Indexing this text can help when searching for keywords, but what if you don’t know the keywords? What if you don’t know where to start?

Unstructured data analysis software allows machines to ‘read’ unstructured text. This could give an understanding of what entities (for example people, organisations or locations) are present in the text. Some of these tools also extract relationships from the sentences and are even multi-lingual making the initial analysis of large amounts of documents quick and easy.

Visual/Audio Analysis

Manually documenting the content of video and audio files is a time-consuming process. Depending on what you are investigating it may also be an impossible task.

The use of cognitive technology to transcribe or even reference known items held within audio and video content can unlock the intelligence held within this kind of media.

Intelligence Data Stores

Having a secure and audited place to hold intelligence as part of a case is of real importance to a forensic accountant. A place where all details of an investigation can be quickly and accurately searched.

A lot of good data stores use the POLE model (People, Objects, Locations, and Events). By framing data in this way, relationships can easily be captured and explored. Intelligence stores will also hold relevant charts, visualisations or even referenced documents, ensuring everything of relevance is in one place.

Companies House Integration

Companies house is a great free resource where you can get access to company information, current and resigned officers or insolvency information.

Making regular queries of companies house through their website is easy, but taking that information and collaborating it with other data sources can be a manual task which is time-consuming. By integrating this resource alongside your intelligence data store it is easier and quicker to see the full picture.

Are you a forensic accountant? Are the technologies above of interest? Or are you struggling with some other form of data you’ve received? Either way, S-branch would love to help.

 

Tackling duplicate data in i2

Smart Matching

Duplicates in data is always an issue. An issue which is magnified once you start trying to load said data into a visual analysis tool like i2. The nature of tools like i2 is that they find connections between entities; in order for that to work the data must be clean a free of duplicates.

In reality, it’s not always feasible or realistic to get rid of all duplicates. For this reason, the i2 analysis suite does provide some matching functionality

Smart Matching – i2 Analyst’s Notebook

To explain it in a simple way: Smart matching in analyst’s notebook is achieved by categorising entities and applying almost human matching logic.

For example, a visualisation may contain a police officer, a victim, a suspect, a prisoner, a male and a female entity. These entities could all be categorised as People. As humans, we could look at a list of people and would instinctively apply matching logic. We’d know that date of birth was important. As are surnames, although they may change when someone marries. We know that Chris could be spelt in several different ways and (as with any data) there may be spelling mistakes.

i2 smart matching can do the same thing. It will apply a different set of logic to different categories of entities to give an almost human matching ability.

Smart matching example

Smart matching example

Great! Problem solved! Well yes, if you have a smallish set of data then smart matching will work well. However, for smart matching to work, the entities need to be on an analyst’s notebook chart. This may start becoming an issue once you hit the 20,000 entities mark! You could try some kind of batching process but it’s likely to be time-consuming and not entirely accurate.

 

Matching in iBase

Matching in iBase example

Matching in iBase example

iBase will not have this problem. Being a database it will have access to all of your records without needing to render them on a visual chart. The problem we have in iBase is it doesn’t have access to the same smart matching logic. With iBase we can match someone with exactly the same surname. Or, the exact same surname and date of birth. While useful, this will miss a lot of duplicates and therefore needs to be used with care.

Custom Matching Functionality

For some of our clients, the functionality described above didn’t meet their needs. They either had too many entities or wanted to find duplicates of completely different entity types, or maybe over different databases.

For these clients s-branch developed an external matching solution, using an existing string matching theory called Levenshtein distance. Levenshtein distance is a string metric for measuring the difference between two sequences. Informally, the Levenshtein distance between two words is the minimum number of single-character edits (insertions, deletions or substitutions) required to change one word into the other. What’s great about using Levenshtein distance is that it gives us a score. Meaning we can see narrow down our results to only the most relevant matches.

Custom matching example

Custom matching example

The external matching is computed outside of iBase. This is done to minimalise the effect on database performance but also allows us to iterate through entire datasets programmatically. This means we can compare duplicates regardless of whether they are the same entity type; or even the in the same database. The output from this is a matching score which is then imported into iBase as a ‘AutoMatch’ link. Our clients can then use existing iBase functionality, such as queries and sets to review the matches found and if applicable merge them through the normal UI.

You can read more on Levenshtein distance here.  To find out more about using Levenshtein distance with iBase or to discuss your own duplicate problem, please contact us