All resources

Data Curation: The Driving Force Behind Successful Language Models in Capital Markets

Trading chat is unstructured, jargon-packed and context-dependent. This brief sets out how Sense Street’s data curation team turns it into the foundation for high-performing models.

Introduction

At Sense Street, we understand the critical role of data curation in fine-tuning language models. Trading chat data, in particular, presents unique challenges due to its highly unstructured, jargon-packed and context-dependent nature. Sense Street’s data curation team combines a deep understanding of financial market dynamics with advanced annotation techniques, providing the necessary foundation for high-performing models tailored to the complexities of trading chats.

Effective data curation rests on three interconnected components: designing annotation frameworks — developing clear guidelines to extract relevant information in a structured format; data quality assurance — ensuring the accuracy and consistency of annotations through meticulous validation; and maintaining adaptive annotation schemas — continuously updating schemas to reflect evolving market trends and data patterns.

1Designing annotation frameworksClear guidelines, structured output2Data quality assuranceAccuracy and consistency, validated3Maintaining adaptive schemasUpdated as markets and data shift
Three interconnected components

Designing annotation frameworks

The first stage in developing an effective language model for a specific asset class is the creation of a well-structured annotation framework — determining what elements within the chats will be annotated and selecting methods for annotating them. These decisions form what we call an annotation schema.

An annotation schema serves two primary purposes. Firstly, it provides comprehensive guidelines for all annotations throughout the project, ensuring consistency and clarity. Secondly, it establishes the direction for the model’s final output. The schema must balance two critical aspects: alignment with the objectives of the end product, and the flexibility to accommodate future modifications or new data formats.

Step 1 · Getting to know the data

Schema creation begins with a deep understanding of the dataset. Data exploration includes domain research — studying market dynamics, traded instruments and negotiation conditions to contextualise the data; applying theoretical knowledge to the data — deciphering trader jargon and abbreviations and documenting them as glossaries; and unravelling implied meaning and conventions — direct collaboration with clients and domain experts to understand the communication conventions unique to a given dataset. That last one is a game-changer for reconstructing the full intended meaning from the chats.

Step 2 · Defining entities

Once a thorough understanding of the dataset is achieved, the next step is defining the entities to be annotated and extracted. Entities are the fundamental building blocks of an annotation schema, representing the elements of the world described in the data. During annotation these entities are labelled, classified into pre-defined categories, and organised into hierarchical structures that reflect the relationships within the text.

Defining entities involves addressing two essential questions: what categories are relevant for the task, and what qualifies as a positive example for each category.

Case 1 · Need for precise definitions of categories

In the RFQ schema, PRICE is one of the crucial categories we want to capture. However, not all prices found in chats are relevant from the perspective of the RFQ — a lunch receipt is not a level. The definition of PRICE for the project needs to specify that only prices relating to negotiations of financial instruments count as positive examples of the category.

Case 2 · Polysemy

Both designing the entities and applying the definitions in annotations requires taking into account the coexistence of multiple meanings for a word or phrase. Depending on context, the same word can represent several different categories we’ve defined. For instance, “Italy” may refer to a holiday destination, a bond issuer, or even a specific bond.

CHATBANKAAPL 2044 still 99.9/RELEVANTCLIENT80$ receipt, these burgersNOT RELEVANTCLIENTTarget is 100.1, can you improve?RELEVANT
Only prices from instrument negotiations count as PRICE

Step 3 · Understanding the importance of granularity

Defining entities requires determining how broad or narrow annotated categories should be. While introducing finer categories increases the complexity of the schema and the number of labels necessary in annotations, it plays a huge role in capturing all the nuances of meaning and maximising the potential of the data.

Case 1 · Annotation of bonds

Bonds, as complex entities consisting of several attributes, are annotated using the method of layering labels. The entire string of attributes is first annotated together as a BOND. Additionally, each granular attribute within the BOND span — issuer, coupon, maturity or ISIN — is assigned to its respective category and annotated separately. This provides the model with structured information about the granular categories that constitute a bond description, enhancing its ability to recognise and link the varying formats used for referencing bonds.

Granular annotations also enable leveraging the full information contained in trading chats by highlighting bond attributes that drive decisions. When a client rejects a bond due to its issuer or maturity, granular labels allow extra information about the client’s interest — or lack of it — to be assigned to the specific bond attribute. This captures not only the status of one RFQ but also broader trading preferences expressed by the client.

ONE SPAN, LAYERED LABELSBMW3.25%2025USU09513JJ95issuercouponmaturityISINANNOTATED TOGETHER ASBONDBOND · 2025 (maturity)BOND · BMW (issuer)EACH ATTRIBUTE ALSO ANNOTATED ALONE
The whole span, then every granular attribute

Step 4 · Leveraging possible methods: descriptive labels

Designing annotation frameworks is about selecting the tools that work best for the task you’re trying to solve. One such method is the use of descriptive labels — free-text descriptions that represent the chain-of-thought of the annotator and emphasise what logical conclusions were drawn from the text. With descriptive labels, the model is provided not only with the correct answer for the task, but also an explanation of how the annotator arrived at it.

The IOI (Indication of Interest) schema captures sentiment towards broader market topics, reflecting client interest or disinterest in trading in specific areas. The challenges of interpreting IOIs include implied dependencies between entities and ambiguous expressions occupying a grey area between IOIs and neutral statements. The strength of IOIs varies significantly, and that nuanced gradation of interest can be effectively reflected through descriptive labels.

Data quality assurance

Following the design of an annotation schema, the next critical stage involves implementing the framework in practice. The primary objective throughout is to maintain consistency and quality of annotations across the entire team. At Sense Street, achieving high-quality annotations relies on three core components.

Expert team

A skilled team is essential for effective data curation. Each data analyst at Sense Street brings a unique combination of proficiency in multiple languages, deep domain knowledge in financial markets, analytical expertise, and sensitivity to the subtleties of human interactions. This diverse skillset allows the team to maximise the value extracted from complex datasets.

Comprehensive documentation

Thorough and accessible project documentation ensures a consistent approach to interpreting nuanced and intricate data. Annotation guidelines act as a critical reference point for the team, enabling precise and consistent annotations.

Review loop

Sense Street’s workflow incorporates a robust review loop. Initial annotations are prepared by data analysts following the established guidelines. Reviewers validate the annotations, providing detailed written feedback if any corrections are needed. Updated annotations are then reviewed again to ensure compliance and accuracy.

The review loop fosters continuous improvement and ensures annotation excellence. Its advantages include effective training — new team members quickly adapt to annotation schemas through iterative feedback; error minimisation — reviews significantly reduce annotation errors, ensuring high precision and adherence to guidelines; pattern recognition — reviewers identify recurring errors, offering valuable insights into schema clarity and documentation quality; and identifying further research areas — continuous exchange within the loop highlights areas requiring further research or schema adaptation.

Maintaining annotation schemas

Annotation schemas are not static — they need to constantly respond to changing market dynamics as well as new formats and edge cases found in the data. This adaptability ensures that schemas remain relevant and effective: new formats feed the annotation schema, which drives annotations, which mean more data analysed, which surfaces new formats again.

NEW FORMATSANNOTATION SCHEMAANNOTATIONSMORE DATA ANALYSEDCONTINUOUS LOOPNEVER STATIC
Schemas respond to changing markets and new formats

Case 1 · Expanding the RFQ schema

The iterative nature of schema updates mirrors the snowball effect — as knowledge about the dataset grows, the schema expands to accommodate new findings. Take the possible number of bonds in an RFQ. At stage one, 1 RFQ = 1 bond, because most RFQs are about buying or selling a single bond. At stage two, 1 RFQ = 1 or 2 bonds, because switches are RFQs that involve an exchange of two bonds. At stage three, with more annotations, edge cases are found — “switch idea for you: I sell 10 mln BMW 24, I buy 5 mln BMW 26 and 5 mln BMW 27” — and 1 RFQ becomes 1, 2 or 3 bonds.

EXPANDING THE RFQ SCHEMASTAGE 11 RFQ = 1 BONDA single bond bought or soldSTAGE 21 RFQ = 1 OR 2 BONDSSwitches exchange two bondsSTAGE 31 RFQ = 1, 2 OR 3 BONDSEdge cases found in annotation
“I sell 10 mln BMW 24, I buy 5 mln BMW 26 and 5 mln BMW 27”

About Sense Street

Sense Street is a generative AI company developing natural language systems for capital markets, to help its participants have better and more efficient conversations. The platform integrates large language models into financial enterprises, improving analytics, automation and observability.