Skip to content
✎ FeedbackNew here?
Open the Credibility board →the live surface this page describes — data, filters and history
How well testedDescriptive onlyit describes what happened; it has never earned the right to rank or selectRecorded: FAILED-AS-FACTOR → DESCRIPTIVE / VETO-ONLY

Credibility — did management do what it said? (Patearn) — Canonical Reference

Renamed 2026-08-19. Called CCI (Concall Credibility Index) until then. The old name collided with the Commodity Channel Index — Lambert's momentum oscillator, which this estate does not compute anywhere. Same measure, same numbers; the URL (?p=cci) is unchanged.

One-line definition: Credibility is a point-in-time management-credibility time-series built from Indian earnings-concall transcripts — extracted measurable guidance is graded against available results and compounded into an as-of credibility_series(symbol, date, level, momentum) per company. Credibility here = Concall Credibility Index (the in-repo program name is "Concall Intelligence"), NOT the Commodity Channel Index — it has nothing to do with Lambert's price/volatility oscillator.


1. What it is

A per-company, point-in-time read on whether a management does what it said it would do. The unit of evidence is the promise: on each earnings call, management makes forward statements ("revenue up double-digit", "capex ₹2,000cr", "debt down by FY-end"). Credibility extracts each one, waits for the resolving result, grades it MET / PARTIAL / MISSED against the actual, and rolls the settled track record into a credibility level (kept-promise hit-rate), a momentum (level vs the prior period), and two symmetric tapes — an EARNING_TRUST tape (a name earning credibility) and a DETERIORATION tape (a name losing it). the desk's framing: credibility is a price, not a stamp — a stock at 500 won't stay 500, and neither does trust; it moves as concalls land and promises settle. Hence a series, not a static score.

The original thesis ( →) was a wedge: a management-credibility signal that FRONT-RUNS price re-ratings — credible managements, guiding up, where the market hasn't priced it. That thesis was return-tested and rejected (see §4). What survives is the honest, defensible thing: a qualitative evidence dossier per name — the promise ledger with verbatim quotes, variances, and outcomes — shown when a human is researching a stock.

2. Our variation vs. the standard technique

There is no textbook "concall credibility index." This is proprietary, and it is deliberately not the two naive things a first attempt would build:

What makes it PIT and honest: every series point is computed using only what was knowable at that concall date — promises resolved by T, quantification of promises made by T, deterioration flags seen by T — with no look-ahead; unproven names (no settled promises) are capped below A rather than flattered; thin samples are shrunk toward a neutral prior; and the whole thing was subjected to a pre-registered falsification gate before any edge claim (§4).

3. How it works (methodology)

The pipeline, in order (runner: the cci pipeline code; free except the extract step):

1. Ingest (free) — the concall bse code discovers primary exchange announcements; the disclosure capture code archives original PDFs and verified content-addressed text. Legacy the concalls code remains for historical compatibility; new pipeline ingestion no longer calls its vendor discovery. Receipt completeness is separate from call-event completeness. 2. Extract (paid, one-time/transcript) — the concall extract code: one Gemini pass per transcript → structured promise/guidance rows. Pre-processed by the cci normalize code (Hinglish + lakh/crore parsing, "fifteen hundred crore" → 1500, commitment classifier, hedge lexicon where ritual fillers score 0). 3. Results — the fundamentals xbrl code refreshes primary filed actuals in the research archive. Existing settlement retains its historical concall_results and deep-actuals compatibility path; the legacy the concall results code vendor fetch is no longer called by this runner. The independent the commitment worker code revalidates archived original result XML into typed evidence and explicit liquidity measures; unsupported inputs remain unavailable. 4. Deep actuals (the depth lever) — the cci deep actuals code: grades old promises against the 24-year fundamentals_history archive (research.db), not just Screener's shallow ~FY2019 table — so a track record reaches back to ~2017. PIT by construction (only report_date ≤ as_of visible). 5. Settle — the concall settle code: resolves OPEN promises against actuals → MET / PARTIAL / MISSED (level / growth% / margin / directional sign-match), multi-year → ONGOING, no look-ahead. 6. Diff — the concall diff code: deterministic consecutive-transcript deterioration tape (walked-back / dropped guidance). 7. Score — the concall scores code: pure-Python, zero-LLM rank + tier. Machine-owner of the scoring constants (W_GA, W_QR, UNPROVEN_CEILING, DETER_PEN_PER, sample gate, shrinkage prior) — do not restate them here; read the code (and the calculations and weights notesg for the weights + numbers doctrine). 8. Series (the spine) — the cci series code: materializes the PIT credibility_series (level + momentum + tape), reusing the scorer's exact constants so the series and the snapshot never diverge. This is the canonical leak-free as-of credibility that every downstream (the gate, the backtest, the Rotation Map, the dossier) reads. ~18,944 PIT points at last full build. In the runner since 2026-09-01. It was previously flagged as a hand-run step outside the chain — and nothing ran it: measured 2026-08-31, credibility_series was 48 days stale while the resources code and the flows code served it as live. Freshness alarm: python -m src.automation.cci_series --check exits 1 when the series is behind the guidance ledger. 9. Rotation map — the cci rrg code: level × momentum quadrant + the credibility-vs-price divergence read into credibility_rrg. In the runner since 2026-09-01 , same reason — it was 63 days stale and read live by the strategist view code. Descriptive only: the divergence is an observation, never a buy or sell (GATE B failed leak-free). 10. Coverage denominator + BSE gap-fill — the concall coverage code rebuilds the expected-call denominator first (free, no network — and it is a correctness ordering, because the filler RANKS off that table and a stale one re-targets symbols already filled), then the concall gapfill code hands the worst-off symbols to the concall bse code.capture_symbol, ranked by missing QUARTERS from the concall coverage code rather than by a cursor. In the runner since 2026-09-02 (W-0167), at 5 symbols on the once-a-day --all pass only — this module is invoked twice per weekday and defaulting it on for both would double the network hit. 🔴 Why it exists: the transcripts we fail to fetch are mostly on the companies' own IR sites (only 176 of the failures are on bseindia.com), which refuse a datacenter IP — and BSE hosts the same documents at a 95% success rate. Proven live: 6 symbols → 112 captured, 0 failed. Capture is free and the source is perishable, so it never waits for budget; it does grow the paid extraction backlog, which is the deliberate trade.

Do not paste or duplicate any of these constants/thresholds into this page — they are code-owned. Link.

What is NOT in the chain (corrected 2026-08-31, W-0113). the concall direction code — the LLM intent-direction classifier — is an operator-run one-time backfill CLI, not a pipeline step. the cci pipeline code has never imported it in any commit of its history; the runner's chain is exactly ingest → results → (bounded) extract → settle → diff → score → series → rotation map (the last two added 2026-09-01). This page previously listed concall_direction twice as a live step, which described work that does not execute.

🔴 A rebuild failure in steps 8–9 RAISES; it does not log-and-return. A stale table that a public route serves as current is precisely the failure a Result=success unit cannot express, so the pipeline exits non-zero and the unit's OnFailure pager fires. Pinned by an automated test, which also exercises the staleness alarm in both directions — an alarm never seen to fire proves nothing.

So what produces direction today? Nothing on a timer does. concall_guidance.direction is written only when an operator runs python -m src.automation.concall_direction --backfill; contrary to that module's own docstring, the concall extract code does not set it inline (the column is never mentioned there). The consumer, the concall signals code._polarity, is built for exactly this: it prefers direction when a row has been backfilled and otherwise falls back to a tightened regex, scoped to capex/expansion only so accounting language ("Consolidated revenue") and deleveraging ("reduce debt") cannot misfire. A NULL direction is therefore a designed state, not a fault — the regex fallback is the live default.

It is deliberately not wired to the weekly timer. It costs one Gemini call_extractor per 20 rows across the whole concall_guidance backlog, with none of the --max-calls bounding that cci_pipeline applies to the free-tier budget, and §4 below records the standing decision not to spend further on Gemini extraction for this corpus. Scheduling it would spend against a decision this page already documents.

4. Status, validation & honesty fence

Credibility is FALSIFIED as a factor. It has NO validated long / short / risk edge. Its only defensible role is a DESCRIPTIVE per-name evidence dossier — never a ranked screen, factor, or leaderboard. This is the headline, and it is load-bearing: representing Credibility as a working signal is a blocking error.

**The return-test (2026-06-25, the cci backtest code): PIT credibility_series (level + momentum) × corporate-action-adjusted forward returns (NSE prev_close chain over bhavcopy_rows), de-marketed cross-sectionally by concall-month cohort; 3,523 proven points (n_resolved ≥ 3) across 377 symbols; 3/6/12m. Two independent reviews (code + methodology) reproduced the result on the VPS** before it was recorded.

Test3m6m12mRead
Level: HIGH−LOW excess−3.1%−5.8%−10.0%high-credibility UNDERPERFORMS (low-cohort t up to +3.6)
Momentum: rising−falling+0.3%+0.8%+1.1%weak; both rising and falling beat flat → "moved" not "rose"
Deterioration veto: P(<−20%) event vs non-event6.8 / 6.9%12.6 / 12.5%14.9 / 15.8%NO downside difference (event marginally better)

Spearman ≈ 0. The inverse level print is fragile: n=377 survivor names, concentrated post-2023 (−17% vs −7% pre-2022) and in high-level megacap mean-reversion — a regime print, not a structural factor. It survives size + valuation neutralization but is heavily survivorship-confounded (the universe is current concall-holders, so only surviving low-credibility names are sampled; the delisted blow-ups the deterioration veto is supposed to catch have been removed by survivorship — the veto is structurally blind to its true target).

**Second method, same verdict — Gate B (2026-06-30, the gate residual alpha research code): the orthogonalized residual-alpha regression fwd_ret ~ cred + ROCE + debt + size + 12-1mom + PEAD (Newey-West HAC), n=1119, reading the leak-free PIT credibility_series as the regressor → cred coef −0.00108, t=−3.71, p<0.001, NEGATIVE (R² 0.028). PASS needs positive + significant → FAIL → merge Credibility into pt14, no standalone book.** Credibility re-prints the already-arbitraged quality factor — it is not incremental alpha over quality + PEAD. (Gate A guidance→return also FAIL/WEAK.)

The content follow-up — and why no content chip ships either. After credibility died, the hypothesis was that the real signal was concall CONTENT (growth-intent: debt_reduction / capex / volume / new_product / expansion), not credibility. A first cross-sectional scan (2026-06-25) looked economically coherent (debt_reduction +2.8% / volume +2.3% / new_product +1.8% / capex +1.5% vs cost_savings −1.0%). **But the placebo harness killed it (2026-07-08,, the concall intent code walk-forward on real concall_dt, 9,461 events): six statement-types passed t_cohort ≥ 2 with same-sign halves (debt_reduction +3.52% t2.43, capex +3.07% t2.70…) — yet against the shuffled-date placebo the largest passing type printed observed +1.90% vs null mean +2.75% / p95 +3.66% (inflation 0.52×, empirical-p 0.925). Random windows of the same covered names drift MORE than post-call windows. The old month-granular tilts are recorded "not reproducible on real dates" — covered-universe beta, not content edge. Guidance therefore stays a candor/promise descriptive axis; no content chip ships as an edge.** Do not overstate the content angle.

Why more extraction cannot fix it (the breadth ceiling). The binding defect is BREADTH / survivorship — the 377-name current-holder universe — NOT corpus depth. No amount of additional Gemini extraction changes the universe you can sample; it only deepens names already in a survivorship-biased set. Therefore the decision (D-recorded): do NOT spend (~₹2,500) to complete the transcript corpus for any factor/screen use. Transcript CAPTURE stays running (free) as a research asset; EXTRACTION stays deferrable (reads on-disk .txt, loses nothing when paused).

Consistent with doctrine: credibility is a veto / context layer, never a ranker or fundable alpha (the ledger's corollary: price strength is the only gross forward-return engine; value/quality/credibility/accumulation are veto/filter/context layers, not rankers). This negative result is a recorded benchmark — reproduce with python -m src.automation.cci_backtest --mode both.

5. Where it lives (code· routes· DB· timers)

Analytics (the automation folder): the cci pipeline code (runner)· the concalls code (ingest)· the concall extract code (Gemini)· the cci normalize code· the concall results code· the cci deep actuals code· the concall settle code· the concall diff code· the concall veto code (forensic gate)· the concall scores code (rank/tier, constant-owner)· the cci series code (PIT series — the spine)· the concall clock code (PIT event dates)· the concall signals code (content/growth-intent materialization — descriptive, see §4)· the cci backtest code (falsification)· the cci rrg code (credibility rotation + ÷price divergence — descriptive map, explicitly not a buy/sell signal).

Operator tools — run by hand, NOT in the chain and NOT on any timer: the concall direction code (LLM intent-direction backfill over concall_guidance.direction; see §3 "What is NOT in the chain"). Listing it above as a pipeline step is the defect corrected under W-0113 — nothing imports it, so nothing runs it unattended.

Falsification gates (the cci folder): the gate guidance return code (A)· the gate residual alpha code (B — the standalone-vs-merge decision)· the gate golden discrimination code· the common code.

Web (the web folder): the credibility fingerprint code → /dash/credibility (the promise-vs-delivery fingerprint flagship; also embedded as the stock dossier's Credibility tab)· the growth view code → /dash/growth (content/growth-intent board — descriptive). Surfaces: /dash/concalls (avoid-tape + credibility-leaders board), /dash/stock Credibility dossier panel, the credibility· cci screener column-group, and Pat NL flows (credibility / deterioration queries). Every surface carries the descriptive-only / no-forward-signal copy.

extension: bounded free reconciliation now follows extraction. It imports missed L1 statements by anti-join, retains provisional identities for specialist review, and reevaluates reviewed cases when evidence changes. It does not alter historical credibility grades. Exact schemas, coverage limits and rollout status: the concall intelligence design notes

DB (SQLite; 9 base tables per): concalls, concall_results, concall_ebitda_watch, concall_guidance (the promise ledger), concall_expectations_vs_actual, concall_behavior (AI-read, not ranked), concall_redflags, concall_scores, concall_coverage (survivorship spine) — schema in the db code SCHEMA_BASE. Plus module-owned tables: credibility_series (the cci series code), credibility_rrg (the cci rrg code), concall_signals (the concall signals code). Legacy transcript text remains under a file on the server<SYM>/*.txt. Verified new capture retains original PDFs in results_daily.db.result_documents, append-only receipt records, and content-addressed text under concalls/objects/. Full historical receipt migration remains open.

Timers (systemd, VPS): hermes-concall-capture.{service,timer} — free full-universe capture, Wed + Sun 09:00 UTC (the scripts folder); hermes-concalls.timer (Mon–Sat 07:00 UTC) drains ~18 Gemini extractions/day oldest-first + settles; hermes-concalls-refresh.timer (Sun) weekly incremental. (Deploy discipline: never run the setup news script or systemctl start a hermes timer mid-day on the VPS — see project deploy notes.)

6. Data & provenance

Source: Indian earnings concall transcripts — Screener "Concalls" section is used only as the index (the concall LIST + transcript URLs); the PDFs download DIRECT from BSE (Referer: bseindia, ~71% BSE / rest NSE / company-IR), parsed with pypdf. Quarterly actuals come from Screener + our own 24-year fundamentals_history (research.db) for deep settlement. This keeps Credibility close to primary sources; the standing wean path off the Screener index for ongoing discovery is the BSE corporate-announcements API (api.bseindia.com/.../AnnGetData/w, probed 200-OK with primed cookies) — a focused adapter, not urgent, consistent with Guardrail #8 (primary sources only; the screener code dependency is the known-remediated exception, not to be extended).

PIT / knowable-at handling (the leak fence): every graded promise is visible only where report_date ≤ as_of; the cci series code composes each point from only what was knowable at that concall date. Entry/knowable clock is two-tier (2026-07-10): the real concall/transcript event date max(concall_dt, transcript_publish_dt) when captured (16,208 rows carry a real concall_dt from the BSE calibration), with period-month-end only as the fallback — so a "Q4 (Mar)" call that actually happened in late-April is not back-dated to Mar-31 (the leak the panel flagged). Gate B's own leak (CL-RES-01) — an early version regressed on the latest composite (embedding resolutions after each anchor) — was fixed by pointing the regressor at the PIT credibility_series; the FAIL verdict is leak-free.

7. Terminology canon

8. Decision & session history

9. Open items / frozen work

10. Sources of truth