Commit graph

1202 commits

Author SHA1 Message Date
30c6fc5747 refactor: improve avatar error handling, storage URL resolution, and upload logging
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 1m46s
2026-09-10 23:05:37 +05:30
c899d7db13 chore: add missing AWS CDK and project dependencies to node_modules
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 2m40s
2026-09-10 22:50:11 +05:30
29f3536a63 Fix indentation in workflow heredoc blocks to prevent YAML scanner error
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 1m55s
2026-09-08 23:28:57 +05:30
5077c04291 Add Supervisor Horizon queue worker and cron scheduler to deployment workflow 2026-09-08 23:23:31 +05:30
f6f6b128f0 Set FILESYSTEM_DISK=public in workflow 2026-09-08 23:16:24 +05:30
bf4490ac45 Add storage:link to setup-ec2.yml workflow 2026-09-08 23:05:42 +05:30
9483685ffd Enable HTTPS with SSL certificate in Nginx and update APP_URL 2026-09-08 22:47:35 +05:30
8dca777f60 Add Facebook OAuth credentials and configure APP_URL in workflow
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 1m19s
2026-09-08 22:39:57 +05:30
bc3e8c3534 Allow storage permissions for both www-data and ubuntu users
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 1m8s
2026-09-08 22:09:34 +05:30
daf9d45c53 Add user:create artisan command
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 1m31s
2026-09-08 22:00:11 +05:30
37772f4c3f Fix duplicate DB values in env generation
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 1m12s
2026-09-08 21:23:00 +05:30
78084b1d58 Install Node.js and build frontend assets with Vite
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 1m41s
2026-09-08 21:19:06 +05:30
ea6cffa474 Update workflow to use the new database credentials for vignesh
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 46s
2026-09-08 21:00:46 +05:30
3391c0b057 Configure PostgreSQL host and run migrations
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 48s
2026-09-08 20:55:56 +05:30
5fec3336c3 Ensure app key is generated if missing
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 50s
2026-09-08 20:52:45 +05:30
d4c5ae2cf2 Fix deployment steps order to prevent permissions issue during composer install
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 50s
2026-09-08 20:49:51 +05:30
8c1a409563 Upgrade PHP to 8.4 and use php8.4-fpm.sock in Nginx config
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 47s
2026-09-08 20:48:09 +05:30
140a1f58c1 Enable debug mode and ensure storage directories exist
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 27s
2026-09-08 20:41:37 +05:30
aedd786edd Add deployment steps to clone code and set up Laravel
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 40s
2026-09-08 20:36:51 +05:30
8732e33e97 Fix YAML syntax error in setup-ec2.yml
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 31s
2026-09-08 20:31:42 +05:30
eb680fd82a Add Nginx proxy setup for socialmedia.tradhox.com 2026-09-08 20:30:30 +05:30
73df5a19b1 Specify valid container image
All checks were successful
Setup EC2 Tools / setup-server (push) Successful in 1m58s
2026-09-08 20:22:59 +05:30
362dab5ed2 Trigger workflow
Some checks failed
Setup EC2 Tools / setup-server (push) Failing after 2s
2026-09-08 20:21:33 +05:30
0d639d3d1b Fix github action URL prefix
Some checks failed
Setup EC2 Tools / setup-server (push) Failing after 3s
2026-09-05 21:08:42 +05:30
38c8b3e419 Fix action uses url to point to github 2026-09-05 21:07:39 +05:30
4c7ecf4bdc Configure workflow to run on push 2026-09-05 21:06:39 +05:30
13151237b4 feat: add Forgejo workflow to automate LEMP stack installation on EC2 2026-09-05 21:02:21 +05:30
ee53672229 refactor: remove legacy agent and cursor configuration files and documentation 2026-09-05 20:56:33 +05:30
9f9ecf400e refactor: remove all Docker configuration and orchestration files 2026-09-05 20:52:37 +05:30
Paulo Castellano
d8149ef058
Remove the automations module (#333)
* Remove the automations module

Drops the visual workflow builder end to end: actions, node runners,
commands, jobs, models, observers, policy, resources, requests, routes,
broadcast channel, scheduler entries, Horizon supervisor, Vue pages and
components, canvas undo/redo history, CodeEditor, translations, factories
and tests.

A new migration rewrites posts.created_via = 'automation' to 'web' and
drops the five automation tables. The CreatedVia::Automation enum case is
removed. The 2026-08-21 social-account identity migration now skips its
automation node repointing when the automations table no longer exists,
so it stays re-runnable after the drop.

Orphaned dependencies removed: simplepie/simplepie, @vue-flow/*,
codemirror and @codemirror/*.

* Drop the orphaned common.beta translation key

* Remove automation leftovers: ResolvableUrl rule, feed fixtures, useShortcut, nav badge

* Drop the unused chart wrapper and @unovis packages

* Address review: keep the shipped migration untouched, drop the dead previewOnly chain

- Restore 2026_08_21 migration to exactly what production ran; the
  rehearsal test now recreates the automations table it expects instead.
- Mark the drop migration's down() irreversible like its siblings.
- Simplify DropAutomationTablesMigrationTest to the sibling shape.
- Remove previewOnly / aiGenerateVariants: the only caller that set the
  prop was the deleted automation Generate node.
- Run pint over lang/*/common.php after the beta key removal.

* Exercise the drop migration against the real automation tables and their FKs

* Drop DuplicateIdentityRehearsalTest: it re-ran a frozen migration that reads the removed automations table
2026-09-05 11:03:23 -03:00
Paulo Castellano
82da3fd64b
Show disconnect alongside reconnect on lost social connections (#332)
When a social account's token expired or the connection was lost, the
Connections card only offered "Reconnect". Users who wanted to drop
that account and connect a different one had no way to do it from the UI.

The lost-connection card now renders both Reconnect and Disconnect,
reusing the existing disconnect confirmation flow. The connected-state
Disconnect button gets a data-testid too so browser tests can target
either state by account id.
2026-09-04 18:29:10 -03:00
Paulo Castellano
c40c740f51
chore(release): v1.0.9 changelog, customer email + thumbnail artifacts (#331)
* chore(release): add v1.0.9 changelog, customer email, and thumbnail artifacts

* chore(release): rewrite v1.0.9 customer email as a webhook announcement

* chore(release): tighten v1.0.9 email around what people asked for

* chore(release): rewrite v1.0.9 email around webhook value

* chore(release): write a plain v1.0.9 webhook announcement email

* chore(release): close v1.0.9 email with a G2 review ask
2026-09-04 15:17:33 -03:00
Paulo Castellano
a932c51bbb
Expose workspace webhooks through the API and MCP (#330)
* Expose workspace webhooks through the API and MCP.

The same create/update/test/rotate/replay/delete flow now lives in Actions so the web UI, REST API, and MCP tools stay in lockstep.

* Keep webhook validation local to each web, API, and MCP entry point.

* Extract MCP webhook rules into request classes and close remaining API/MCP review gaps.

* Tighten webhook updates to a field whitelist and reset failures only on re-enable.

* Treat a mismatched webhook replay as not found and mark secret rotation destructive.
2026-09-04 12:00:32 -03:00
Paulo Castellano
91a1d5ca1e
Slim webhook log broadcasts to fix Reverb payload-too-large (#329)
* Stop broadcasting webhook payloads so Reverb does not reject live log updates.

The Echo event only needs list metadata; Inertia reloads the full log after the status lands.

* Type webhook Echo payloads separately and cover delivered, failed, and pending broadcasts.

The show page still receives payload and body over HTTP; Echo only carries list metadata.
2026-09-04 10:04:07 -03:00
Paulo Castellano
58d8e066b5
Add workspace webhooks and drop the unused automation webhook node (#326)
* Add workspace webhooks and drop the unused automation webhook node.

Give workspaces HMAC-signed outgoing webhooks for the post lifecycle, with retry, auto-pause, replay, and live logs, and keep HTTP Request as the only outbound automation node.

* Tighten webhook controller and validation after review.

Drop the redundant workspace redirects, prune logs without counting, and validate events/status with Rule::enum.

* Move leftover webhook UI copy behind i18n.

HTTP status phrases, delete-cancel, and validation attribute names were still English literals.

* Build the webhook-paused email through Maizzle.

The hand-written Blade skipped the shared layout, header, and footer used by the other mail templates.

* Cover real webhook dispatch paths and restyle the webhook pages.

* Ask for the shared delete keyword when confirming a webhook delete.

The endpoint URL is a poor confirm string; posts and assets already use the common "delete" keyword.

* Fix webhook review blockers so CI can go green.

Drop leftover French automation keys, stop mutating Inertia log props, and show delivered_at instead of created_at.

* Close the remaining webhook review gaps.

Keep Echo log updates across infinite scroll, align the channel with the policy, persist log ids across retries, and fail unknown automation nodes without throwing.

* Stop webhook delivery after disable and record last sent only on success.

Queued jobs now skip paused or disabled endpoints unless the user replays, and changing the URL re-pings it first.

* Limit webhooks to owners and admins, and encrypt signing secrets.

Members can no longer create or inspect outgoing integrations, and secrets stay encrypted at rest.

* Cover webhook secret hiding, skip-ping, and failed-delivery edges.

* Send the full post on webhooks after labels and platforms are saved.

* Fix webhook payloads for integer media ids and type webhook status.

* Split the webhook show page into focused components.

* Reset live webhook logs when switching endpoints.

* Keep the newest webhook logs at the top after live merges.

* Cast media item ids to string without the extra scalar check.

* Add post.unscheduled webhooks and put the log id on the envelope.

Unscheduling is now a first-class event, and receivers can send the delivery id back so we can find the matching log.

* Translate webhook event names in the UI.

* Make the webhook show page full-width and stop stacking flash toasts.

* Translate remaining webhook UI copy in every locale.

* Sign webhook pings and drop author email from the payload.

* Send signed webhook tests after create instead of pinging on save.

Create and update only block private URLs so the receiver can copy the secret first. The show page then sends a signed webhook.test with an object data envelope.

* Polish webhook test UX and always mint the dispatch log id in the job.

Keep send-test in the actions menu (its own group) and drop the leftover constructor param so retries reuse the serialized id instead of a caller-supplied one.
2026-09-04 09:43:29 -03:00
Maciej Dzierżek
ca497e1f3c
perf(horizon): skip queues for platforms that are disabled (#321)
* perf(horizon): skip queues for platforms that are disabled

The social-publishing supervisor listens on Platform::allQueues(), which maps
over every enum case regardless of the per-platform *_ENABLED toggles. With
minProcesses => 1 that means one worker per supported platform - even on an
installation that only ever connects two or three of them.

On a single-workspace self-host that was 19 Horizon workers at roughly 75 MB
each; filtering by the toggles brought it to 9 and cut container memory from
1.66 GB to 1.06 GB, with no change in publishing behaviour.

Note on the implementation: the filter reads env() directly rather than calling
Platform::isEnabled(). Config files load alphabetically, so config('trypost.*')
does not exist yet while horizon.php is evaluated - isEnabled() would silently
return its default of true and the filter would be a no-op. This bites at
config:cache time, so it is invisible in tinker.

* refactor(horizon): filter disabled platform queues via enum

Move queue filtering to Platform::enabledQueues() and apply it from
AppServiceProvider after config has loaded, avoiding env() parsing in
horizon.php while keeping isEnabled() as the single source of truth.

* refactor(horizon): use enabledQueues directly in horizon config

Remove AppServiceProvider boot override and let isEnabled() fall back
to env when trypost config is not loaded yet.

* chore: enable Eloquent strict mode in all environments

* refactor(platform): replace enabled env key derivation with explicit match

* refactor(platform): simplify isEnabled using config default fallback

* test(platform): cover publishing queues and enabled toggles exhaustively

* refactor(platform): collapse isEnabled env fallback into one method

* refactor(platform): drop filter_var and rely on env boolean casting

* revert: keep Eloquent strict mode out of production

shouldBeStrict() in production would throw on lazy loads in queued
publish jobs and can stop posting. Restore the Laravel default.

---------

Co-authored-by: Paulo Castellano <paulo@castellanos.llc>
2026-09-03 13:51:03 -03:00
Paulo Castellano
abe687d0fa
Fall back to display name when a published-post email has no username (#325)
* Fall back to display name when a published-post email has no username.

Facebook Pages often have no vanity username, which left the email as "Facebook Page (@)". Use username, then display name, and omit the empty parentheses when both are missing.

* Move the notification account label onto PostPlatform.

The fallback belongs on the row that always has a platform, so mail and in-app notify can call one method instead of guarding a missing social account at every call site.

* Cover the in-app published/failed notification body for Facebook Pages.

The email tests were not enough — SendNotification body is what the in-app inbox shows, and that path also goes through notificationLabel().
2026-09-03 12:51:44 -03:00
Grigory Frolov
0db9e0706d
Visible Terms and Privacy links on the auth screens (#317)
* feat: visible Terms/Privacy links on the auth screens

Social-platform app reviews (TikTok explicitly) require both links to be
reachable from the public site without logging in or opening a menu.

* Cover the legal links with tests and tidy the layout

The links are a compliance artifact an outside reviewer checks, but
nothing asserted they exist. A refactor of this layout could drop the
footer silently and the next platform submission would fail the same
check that prompted the PR. A feature test asserts the shared prop
reaches both guest screens and carries whatever the install configured,
and a browser test asserts the two links actually render to a logged-out
visitor.

Declares legal on SharedData, so the props read through the interface
rather than its index signature and the inline cast goes away.

Adds rel="noopener noreferrer" to both anchors, matching every other
target="_blank" in the codebase, and orders the imports the way eslint
expects so the file lands clean rather than relying on --fix in CI.

Drops the comment explaining why the links are there: that rationale
belongs in the commit and the pull request, and the sibling comments in
this file describe markup rather than justify decisions.

* Reuse the legal sentence the register screen already had

The register screen has shown "By continuing, you agree to our Terms of
Service and Privacy Policy" in production for a long time, translated
into all sixteen locales. Only the login screen was missing it, and the
two URLs were hardcoded inside the translated string, so a self-hosted
install could not point them at its own documents.

So this keeps what already worked and changes only those two things. The
markup moves into one component, which the login screen now renders as
well. The translated sentence keeps its wording and its link labels; only
the href becomes an i18n placeholder that the component fills from
config. That is one line per locale, and no new translation keys.

Reverting the footer out of AuthSplitLayout also stops the links from
appearing on the workspace index and create screens, which reach that
layout too and are seen after login rather than before it.

The sentence no longer hides on a self-hosted install. It was hidden
because it named TryPost's own documents; now that the URLs are
configuration, an install that sets them wants it shown.

A feature test covers the shared prop on both screens and the placeholder
in every locale; a browser test covers the rendered sentence, since the
links are a compliance artifact an outside reviewer checks and nothing
guarded them.

* Let the browser assertions do their own waiting

The test hand-rolled a polling loop in injected JavaScript to wait for
the element to mount, because the project notes say browser assertions do
not wait for SPA paint.

They do. visit() returns a PendingAwaitablePage backed by
AwaitableWebpage, whose __call wraps every method in
Execution::waitForExpectation and retries until the Playwright timeout,
which defaults to five seconds. The plugin even deprecates waitForText in
favour of assertSee for this reason. The loop was re-implementing the
retry that already surrounded each call, less well and with a helper
whose name has to be unique across the whole suite because these are
global functions.

Halves the file and drops the injected script. assertSeeLink also says
more than the old check did: it asserts the labels are links, not just
text that happens to appear.

---------

Co-authored-by: Paulo Castellano <paulo@castellanos.llc>
2026-09-01 18:20:38 -03:00
Paulo Castellano
b689e3902c
Send the ad click IDs to Stripe alongside the UTM parameters (#323)
Signup already captures a click ID per ad network -- gclid, fbclid,
li_fat_id, ttclid, rdt_cid, epik -- but only the UTM parameters reached
the subscription. UTMs say which campaign a customer came from; the click
ID identifies the individual click, which is what Google and Meta need to
match a subscription back to the ad that produced it.

Every key is spelled out at the call site, so what reaches Stripe is
readable in one place without following a constant.

Click IDs are text columns on purpose, since truncating them at 255 would
destroy the value. Stripe caps a metadata value at 500 characters and
rejects the request rather than truncating, so an unbounded column
reaching it would fail the whole checkout and the customer could not
subscribe at all. Values are cut to 500 before they are sent. The UTM
columns are varchar(255) and cannot reach it; this exists for the text
ones.
2026-08-31 23:01:37 -03:00
Paulo Castellano
c1af03308a
Carry signup attribution into the Stripe subscription (#322)
Every account arrives with the UTM parameters that brought it and the
onboarding answers it gave, but none of that reached Stripe. Revenue
lived in one system and the campaign that produced it in another, so
answering "which campaign is paying for itself" meant joining the two by
hand on email.

The account owner's five UTM parameters, persona, goals and referral
source now ride along as subscription metadata. Cashier's withMetadata()
puts them in subscription_data.metadata during Checkout, so they land on
the Stripe Subscription rather than the session: the session is gone in
24 hours, the subscription carries the attribution for as long as the
customer does, and it shows up on the subscription in the dashboard.

Stripe metadata holds strings, so the two enum-cast columns are unwrapped
to their backing value and goals -- a JSON column, the one list answer --
is joined with commas. Passing either through as-is fails the request.
array_filter drops the keys the user never filled in, since an absent key
reads the same in reporting and Stripe treats null as a delete.

The account may have no owner, so every read is null-safe and an account
without attribution simply sends no metadata.
2026-08-31 22:31:09 -03:00
Paulo Castellano
6496588bbc
Defuse links in X posts to avoid the link-post fee (#308)
X bills a post containing a URL at a much higher rate than a plain post, and
its algorithm demotes link posts. The X version of a post now rewrites every
URL non-clickable (https://example.com/post becomes example(.)com/post):
scheme and www. dropped, every dot of the host replaced with (.).

Leaving a single dot intact would still leave a resolvable domain for X to
detect, so all of them are broken. A scheme or www. proves a token is a URL on
its own; a bare host only counts when its last label is a delegated TLD, which
is the one thing telling acme.com apart from Node.js. That check runs against
App\Support\LinkTlds, generated from the whole IANA root zone in every form a
TLD can appear in a post -- ASCII, punycode and the Unicode it decodes to --
because whatever X links is what X bills, so a hand-picked subset would leave
us paying for its gaps. If the regex engine bails out on pathological input the
original content is returned instead of crashing the publisher.

The transform lives in the Platform::X arm of ContentSanitizer, so it reaches
publishing and the app/API/MCP previews from one place and cannot touch any
other network. Off by default; opt in with X_DEFUSE_LINKS.

The editor counts characters and renders its preview client-side and cannot ask
the server on every keystroke, so the rewrite is mirrored in TypeScript. PHP
stays the source of truth: a parity test fails if the two TLD sets drift, and a
browser test drives the real editor so the mirror is covered rather than
assumed. Without it the composer promised text the network never receives.

Character limits now measure the text a reader will see: sanitized, then with
markup resolved away. Measuring the raw draft blocked saving posts that publish
fine and let through posts the network rejects, and counted the editor's HTML
toward the limit. Measuring the sanitized form alone would have counted
Telegram's escaped entities, rejecting messages Telegram accepts.

Empty content is handled once inside the sanitizer instead of by a guard
repeated at every call site.
2026-08-29 15:31:05 -03:00
Jamie Ontiveros
02425038aa
Make the test suite pass on MySQL (#307)
* Give the foreign key a backing index before dropping the unique

social_accounts.workspace_id carries a foreign key, and the composite
unique index is the only one covering it, as its leftmost prefix. MySQL
refuses to drop the sole index backing a foreign key (SQLSTATE[HY000]
1553), so both rehearsal suites failed in beforeEach and never ran a
single assertion on MySQL. Add a plain index on workspace_id first;
PostgreSQL has no such requirement and simply carries it.

This unmasks one assertion underneath that had never executed: the
automation graph comparison at DuplicateIdentityMigrationTest.php:419
depended on JSON object key order, which MySQL normalises on storage.

(cherry picked from commit 98a494bd2205e873321a18232f63b358ae259fdf)

* Compare JSON payloads without depending on key order

MySQL normalises JSON object keys (length, then lexicographic) on
storage, so an identity comparison against a literal asserts how the
driver chose to lay the object out rather than what it contains.
PostgreSQL preserves insertion order, which is why these passed there.

toEqual compares associative arrays recursively without regard to key
order. Applied to every assertion in this class, including the few that
pass today only because their keys already happen to match MySQL's
ordering.

(cherry picked from commit 3124023c548d6c2b8b52126afc6fc5f38d461ea6)

* Match logged SQL without depending on identifier quoting

Four DB::listen predicates matched 'select * from "post_platforms"'.
PostgreSQL quotes identifiers with double quotes and MySQL with
backticks, so on MySQL the predicates never matched, the simulated
mid-run pause never fired, and the race these tests exist to cover went
unexercised while the tests still reported failures elsewhere.

Compare against the unquoted form via a small helper.

(cherry picked from commit 67a81df5de155e80227df748b34cd8b3cfd744f9)

* Cast raw boolean reads in tests so they pass on MySQL

Three assertions read oauth_refresh_tokens.revoked through the query
builder rather than Eloquent, so no cast applies and the driver's native
representation leaks into the test: a real boolean on PostgreSQL, 1 on
MySQL. Cast explicitly at the call site.

(cherry picked from commit 2911c5c48cf65d24a34a41e667335c40005839a7)

* Use a scheduling date inside MySQL's TIMESTAMP range

MySQL TIMESTAMP columns end at 2038-01-19, so the 2099 sentinel these
tests used is rejected outright with SQLSTATE[22007]. 2037-12-31 still
reads as a far-future schedule and works on both engines.

(cherry picked from commit bde33eb239cdbd3a5567d4c21e1d85302913cdd7)

* Remove the duplicate-identity migration scenario test

The suite rebuilt a pre-migration schema by dropping the unique index in
beforeEach and re-running the migration by hand, exercising a database
state the application never runs in.

* Fix the MySQL rollback path and run CI on both engines

The migration's down() dropped a unique whose leftmost prefix is an FK
column, which MySQL refuses when nothing else backs the constraint
(SQLSTATE 1553). It now creates a standalone index first, so
migrate:rollback works on MySQL and stays a no-op change for PostgreSQL.
up() is untouched: every database already migrated keeps its schema.

The rehearsal test calls that down() instead of hand-rolling the drop,
so it exercises the real rollback rather than an imitation of it.

Matches logged SQL through the connection's query grammar rather than
stripping quote characters, and adds a MySQL leg to the backend CI job.

* Use a readiness check both database images can run

mysql:8.4 installs mysql-community-server-minimal, which ships neither
mysqladmin nor the mysql client, so a mysqladmin health command never
succeeds and the service never reports healthy. Both images run their
init phase without networking, so an open port is the point either
engine starts accepting connections - one check covers both, and the
per-engine matrix key goes away.

* Use each engine's own readiness tool

pg_isready and mysqladmin ping are what the respective images ship for
this, and the mysql image's entrypoint invokes mysqladmin itself, so it
is present. Keeps 20 retries, which MySQL needs to finish initialising.

* State the two-engine ceiling as a rule, not a test detail

The 2038 TIMESTAMP limit binds anything written to the column, not just
the sentinel dates in fixtures, and the same reasoning generalises: what
the app supports is the intersection of both engines.

* Let the release image connect to MySQL

The published image installed only pdo_pgsql, so DB_CONNECTION=mysql
failed with "could not find driver" before any query ran - the app
supports MySQL but the image people actually deploy could not reach it.
mysql-client mirrors the postgresql-client already present, for
artisan db and dumps.

* Keep "backend" a single required status check

Matrixing the job split its check in two, so the "backend" context the
branch protection requires was never reported and every PR sat waiting
on it. The matrix is now "tests" and a small "backend" job gates on it,
which keeps the required check stable however many engines the matrix
grows to - and leaves the open PRs mergeable without a rebase.

---------

Co-authored-by: Paulo Castellano <paulo@castellanos.llc>
2026-08-29 11:05:33 -03:00
Paulo Castellano
c66c68b115
chore(release): v1.0.8 changelog, customer email + thumbnail artifacts, and /release copy rules (#305)
* chore(release): add v1.0.8 changelog, customer email, and thumbnail artifacts

* chore(release): sharpen the v1.0.8 email and encode the rules in /release

The subject and section headers named areas (Facebook Pages, Asset
Library) instead of what got better, so a subscriber could not tell
from the inbox what the release did for them.

- Lead on the widest-reach change (Business Portfolio Pages) rather
  than the newest one, and match the section order to the subject
- Say plainly that Asset Library reuse is new for the API and MCP
  only, since the in-app path already worked
- Drop the four bullets that repeated a theme section verbatim
- Reorder the thumbnail chips to follow the sections

.claude/commands/release.md carries the same rules so the next release
starts there: subject pattern, ordering by reach, the already-existed-
in-app trap, header wording, and no duplication between prose and lists.
2026-08-29 09:36:28 -03:00
Hafiz Muhammad Moaz
06fd0ba53d
docs: surface per-provider AI model override env vars (#290)
* docs: surface per-provider AI model override env vars

.env.example only showed AI_TEXT_PROVIDER/AI_IMAGE_PROVIDER/
AI_AUDIO_PROVIDER (which provider to use), never the per-provider
model overrides (OPENAI_IMAGE_MODEL, GEMINI_AUDIO_MODEL, etc.) that
config/ai.php already wires up for every provider capability. A
self-hoster reading only .env.example had no way to discover these
existed.

These aren't new: laravel/ai already resolves each capability's
model independently per provider (confirmed in vendor - e.g.
OpenAiProvider::defaultImageModel() falls back to a hardcoded image
model, never to the text model), so there was never actually a
text-model fallback bug to fix in code. The gap was purely that the
override knobs were invisible outside of reading config/ai.php's
source directly.

* docs: document AI provider keys, URLs and example model values

The override block landed with bare variable names and only covered the
five providers whose API keys were already in .env.example. Filled in
what a self-hoster still could not discover:

- Example values on every model override, taken from laravel/ai's
  current per-capability defaults, so the ID format is visible (an
  Ollama tag, an OpenRouter vendor/model path and a bare OpenAI ID
  look nothing alike).
- The keys and URLs for the providers .env.example already advertises
  as valid but never documented: XAI, GROQ, MISTRAL, DEEPSEEK,
  OLLAMA_URL and the openai-compatible endpoint. Ollama was the
  scenario in the original report and OLLAMA_URL matters more there
  than OLLAMA_TEXT_MODEL does.
- GROQ/MISTRAL/DEEPSEEK text models, missing while their providers
  were listed as options.
- openai-compatible has no built-in default model and throws without
  OPENAI_COMPATIBLE_TEXT_MODEL, so that one is flagged as required.
- Corrected the audio provider list: gemini and openrouter implement
  AudioProvider too. The image and audio lists are no longer written
  as exhaustive.
- Mirrored all of it into docker/.env.docker.example, which carried the
  same block untouched and is the file Docker self-hosters copy. Its
  OLLAMA_URL points at the host instead of localhost.

---------

Co-authored-by: Paulo Castellano <paulo@castellanos.llc>
2026-08-26 11:13:06 -03:00
Jamie Ontiveros
6e10e394a9
Use whereLike for search so MySQL works alongside PostgreSQL (#302)
Seven search call sites used the `ilike` operator, which only PostgreSQL
understands. On MySQL they raise a syntax error, so post, asset, label,
signature and workspace-member search — plus the MCP list-posts tool —
were unusable on an engine `config/database.php` has always supported and
the docs advertise.

Replace them with `whereLike($column, $value)`, which the query grammars
translate per driver: PostgresGrammar emits `ilike` and MySqlGrammar emits
`like`. The generated SQL on PostgreSQL is therefore unchanged.

Verified by running the full suite on both engines:

  PostgreSQL 16    3888 passed, 0 failed
  MySQL 8.0.46     one pre-existing failure fixed, none introduced

Also adds case-insensitivity assertions to the five affected suites that
lacked them, and search coverage for ListPostsTool, which had none.

Note for MySQL installs: `like` is case-insensitive by virtue of the
column collation, not the operator. Under the default `utf8mb4_unicode_ci`
it is also accent-insensitive, so a search for "cafe" matches a stored
"café" — PostgreSQL's `ilike` does not. That difference comes from the
collation rather than this change. A `_bin` or `_cs` collation would make
search case-sensitive on both.

Co-authored-by: Paulo Castellano <paulo@castellanos.llc>
2026-08-26 10:59:31 -03:00
Paulo Castellano
02e44b9785
Connect the Pages a login only reaches through a Business Portfolio (#301)
* fix: Facebook Page fetch missing New Pages Experience pages

/me/accounts silently omits Pages that live under Meta's newer
"New Pages Experience" / Business Portfolio model, even when the
token's granular scopes show the Page was explicitly granted -
confirmed via Meta's own Access Token Debugger against a live
account whose Page returned zero results from /me/accounts but
resolved fine when queried directly by ID.

Falls back through Business Manager's owned_pages/client_pages
(via the existing business_management scope) when /me/accounts
comes back empty, so Pages under that model are still found.

* fix: Instagram-via-Facebook has the same New Pages Experience gap

Same root cause and fix as the FacebookController fetchPages()
fallback - the Page/Instagram-linked-Page lookup goes through the
same /me/accounts call and is subject to the same Meta-side gap.

* refactor: one place finds every Page a Meta login can publish to

Both controllers walked /me/accounts and, when it came back empty, the
Business Portfolio edges behind it. ManagedPages now owns that walk for
the Facebook and Instagram-via-Facebook flows alike.

Three behaviour changes come with it:

The portfolio edges are read on every connect, not only when /me/accounts
is empty, and merged by Page id. Someone holding one Page by a classic
role and the rest through a portfolio was auto-connected to that single
Page and never offered the others.

/me/businesses is read only once /me/permissions confirms the login
granted business_management, and a failure anywhere along the portfolio
walk leaves the /me/accounts list standing. It used to escape into the
callback's catch, turning "no pages" into "could not connect" for every
login without the scope.

Pages the login cannot get an access_token for are dropped. Connecting
one produces an account that cannot publish.

* test: pin the Page a login only reaches through a portfolio

Covers the merge with /me/accounts, the access_token filter, the
business_management gate, and a failing portfolio edge leaving the
/me/accounts list intact — plus the connect flow end to end on both
Meta platforms.

The Instagram-via-Facebook request count moves from five to six for
the /me/permissions check.

* refactor: read the portfolio edges without asking permission first

The /me/permissions check saved a rejected /me/businesses call on logins
without business_management, at the cost of running a path nobody had
verified against a live account. The edges already fail soft, so the
check bought log tidiness and nothing else.

* test: cover the portfolio walk's remaining shapes

The multi-Page selection flow behind a portfolio — the case #292 asked a
maintainer to check — plus pages spread across two portfolios, a
paginated edge, a portfolio entry with no id, and a portfolio page
merging with one /me/accounts already returned.

* test: stop the Meta connect tests from calling Graph for real

Http::fake only stubs the URLs it is given; anything else goes out over
the network. These files stubbed /me/accounts and left /me — and now
/me/businesses — unstubbed, so the suite was issuing live requests to
graph.facebook.com on every run. They came back 400 and the code under
test swallowed them, so nothing ever went red while the assertions were
measuring Meta's answer instead of the fixture's.

Every Graph call the connect flow makes is stubbed now, and the files
prevent stray requests so a missing one fails loudly. Inertia's SSR
endpoint is allowed through; it is not what these tests are about.

* fix: tell a denied portfolio edge apart from a throttled one

GraphPaginator throws so no caller reads a failed fetch as an empty list
and auto-connects whatever arrived first. Swallowing that exception on
the portfolio edges gave the invariant away: a 429 on owned_pages left
the merged list holding only the /me/accounts page, and the callback
connected it with no picker.

The exception now carries whether the failure was transient, classified
by GraphError, which already owns Meta's rate-limit and transient code
table. A denied permission reads as "this login reaches no portfolio
pages"; a throttle, a 5xx or a truncated walk is raised.

The walk also stops at MAX_PORTFOLIOS and logs what it skipped. Each
portfolio costs two more paginated edges inside a synchronous OAuth
callback, and nothing bounded that loop.

* fix: three ways the portfolio walk misread what Meta returned

Concurrent Instagram lookups. The Instagram description ran one request
per Page, in sequence, at a 15s timeout each. That list used to be the
Pages someone holds a role on — a handful. It is now the union with
every portfolio's owned_pages and client_pages, so a portfolio holding
hundreds of Pages serialised the OAuth callback past any gateway
timeout, for exactly the accounts the portfolio walk exists to reach.
Meta's ids= batching is no help: each Page carries its own access token
and one call takes one token. The lookups run in concurrent rounds.

A Page without a token is not a Page you don't have. Meta lets someone
decline pages_read_engagement on its per-permission toggles and still
lists the Page, without an access_token. Dropping it inside the walk
left the caller saying "no Pages found, you need to be an admin of at
least one" to an admin. ManagedPages returns everything Meta listed and
publishable() separates what can be posted to, so the callbacks can
tell the two apart and say which one happened.

Stored scopes are what Meta granted. The scope list was written to the
account's scopes column straight from the request, claiming access the
login may have refused — business_management above all, which needs
Advanced Access and is declined by default without it. It now comes
from /me/permissions, falling back to the request when Meta cannot be
asked.

* fix: only drop a scope Meta says was refused

PublishToSocialPlatform::failForMissingScopes() blocks a post when a
platform's required publish scope is absent from the account's scopes
column, so writing that column from /me/permissions can dead-end an
account. Meta does not document that the endpoint echoes scope strings
verbatim, and the edge is paginated, so a scope it never mentions is
unknown rather than refused and stays. Only declined and expired drop.

* fix: keep the portfolio walk honest and cheap

Raise instead of truncating. The ceiling logged a warning and returned
whatever fit, which is the one thing this module refuses to do
everywhere else: if the walk cannot finish, the real list is unknown,
and a truncated list holding exactly one Page would have been
auto-connected without ever showing the picker. It now raises, and the
ceiling rises to GraphPaginator::MAX_PAGES' 100 since the walk no
longer pays for it serially.

Read the edges concurrently. Up to two paginated edges per portfolio ran
back to back inside the OAuth callback. They run in rounds now; a URL
that does not come back cleanly still goes through GraphPaginator, which
owns the single place that logs a Graph failure and decides whether it
is a rejection or an unknown.

Prefer the record that carries a token. Merging kept whichever copy of a
Page id arrived first, and /me/accounts always arrives first — so a Page
listed there without a token buried the portfolio copy that had one, and
the login was told its permission was missing for a Page it could reach.

Describe only the Instagram accounts that survive. The per-Page lookup
ran before filterConnectableIdentities discarded them, spending a
BUC-rate-limited call on every Page only to throw the answer away. The
filter reads instagram_business_account.id straight off the raw Page, so
it needs no lookup to run first.

Two Instagram tests mocked Socialite without usingGraphVersion, so the
callback threw, the generic catch answered, and asserting only
success=false passed on the error path instead of the one under test.

* test: pin that pages survive past the first pooled round

The concurrency test drove 30 portfolios with every edge empty, so the
merge across rounds was never exercised with data in it.

* fix: keep the paging-host guard on the pooled edge walk

GraphPaginator refuses to follow a paging.next that points off the host
the walk started from, so a tampered response cannot carry the access
token somewhere else. Reading the first page of each edge through the
pool and handing its paging.next straight back to GraphPaginator made
that URL the *start* of a new walk, which is the one URL the guard
trusts implicitly — so the first hop went unchecked.

The host is compared before the hand-off now, and a mismatch re-walks
the edge from the beginning so GraphPaginator's own guard is what
refuses it, with its logging.

* fix: stop a cut-short walk from passing for a complete one

optional() was written for "this edge may be forbidden" and answers a
rejection with an empty list. Following paging.next through it gave a
rejected cursor the same answer: page one of a 250-Page portfolio came
back and the rest was dropped, and a single connectable Page in that
fragment would have been auto-connected with no picker. The same hole
sat on /me/businesses, where GraphPaginator is all-or-nothing — a
failure on page two threw away the portfolios page one had already
listed, degrading the connect back to /me/accounts alone in silence.

The exception now carries how many pages arrived. Only a rejection on
the very first request reads as "this edge is not readable"; anything
after that is a fragment and raises. Cursors skip optional() entirely.

A login Meta reports as refusing business_management also stops walking
the edges at all. The controllers already read /me/permissions for the
scopes column, so the answer costs nothing, and the walk was otherwise
spending a request on a certain 403 — and logging it at error level —
on every successful connect by such a login.

An Instagram account with an empty Name connected as display_name null:
data_get's default only fires on an absent key, and describeRound always
writes the key.

* test: pin reconnecting a card only the portfolio still reaches

A card whose Page moved behind a portfolio is the reconnect shape of the
bug this branch fixes, and nothing covered it: the walk has to find the
Page, and filterConnectableIdentities has to keep the original card
rather than offering the portfolio's other Pages.

* fix: an unreadable portfolio must not deny the pages that were readable

The portfolio edges are additive, but every failure in them was raised
and the callback's generic catch turned it into "error connecting" — so
one throttled edge among sixty denied a login the Pages /me/accounts had
already returned, and each retry burned more of the quota that caused it.
Only /me/accounts failing is fatal now; everything else marks the walk
incomplete and keeps what arrived.

What the raise was protecting is kept where it belongs: a lone Page is
only taken without asking when the walk saw everything, or when a
reconnect has already pinned which Page is wanted. Otherwise the picker
opens, and the login can see for itself that its Page is not there.

The ceiling stops pretending. It compared the count after walking every
page of /me/businesses — up to ten thousand ids — so the runaway it
existed to bound had already happened. One request, one page, and more
portfolios than that is an incomplete walk rather than a failed one.

A pooled edge that fails is classified where it lands instead of being
re-fetched, halving the cost of the common client_pages rejection, and
GraphPaginator logs a confirmed rejection at warning: it is Meta
answering the question, not something going wrong.

A login that declined the permission its platform needs to publish is
refused at connect. Meta issues a Page token off pages_show_list, so
declining pages_manage_posts still produced a green account whose every
scheduled post was then hard-failed by failForMissingScopes.

Test fakes address Graph through the config rather than a literal host,
which is what CLAUDE.md asks for and what the newer tests already did.

* fix: an incomplete walk must not answer as if it were sure

Marking the walk incomplete stopped the auto-connect, but every dead end
after it still gave a definitive answer. A login whose only Pages sit
behind a throttled portfolio was told "no Facebook Pages found, you need
to be an admin of at least one" — the exact sentence this branch exists
to stop showing to admins, now arriving for a different reason. The
already-connected and missing-permission answers were equally sure of
themselves.

When the walk could not see everything and there is nothing to offer,
it says so and asks for a retry.

* fix: stop every Inertia test from calling an SSR server

inertia.ssr.enabled defaulted to true and nothing in phpunit.xml turned
it off, so every test rendering an Inertia page issued a real request to
the SSR endpoint. The project does not use SSR, so those calls only ever
failed and fell back to client rendering — quietly, on every run.

Defaulting it off is what the project already assumed, and it retires
the allowStrayRequests hole the Meta connect tests were carrying to work
around it. INERTIA_SSR_ENABLED still turns it back on.

* fix: a taken slot is a fact, not a guess about the listing

Routing every short listing to "try again in a moment" swallowed
network_taken: a workspace that already holds its one Facebook account
was told to retry, forever, whenever a portfolio edge was throttled. That
answer comes from our own rows and does not depend on how far the walk
got. all_connected and page_not_found do, and still yield.

Also: an off-host cursor now stops the edge instead of re-reading page
one, which cost a request and could follow an on-host cursor on the
retry, quietly undoing the guard. Cursor follow-ups are budgeted, since
they cannot be pooled and were the one unbounded serial path left. The
exception's fetched count lost its last reader two commits ago and is
gone. Comments trimmed throughout.

* fix: bound the cursor walk by requests, and keep what it read

MAX_CONTINUATIONS counted edges, not requests: each one then handed off
to GraphPaginator, which follows up to a hundred more pages by itself.
The budget the docblock promised was fifty times larger than it claimed.
Cursors are now followed one budgeted request at a time, so the count
means what it says, and pages already read survive a cut-off instead of
being thrown away with the exception.

A refused /me/businesses is no longer read as "this login has no
portfolios". For a single edge a rejection answers the question; for the
index of edges it means we could not look — and answering complete there
auto-connected the one /me/accounts page while hiding every portfolio
Page, which is this branch's own bug wearing a different hat.

SSR goes back to its shipped default. Turning it off in config to quiet
the test suite would have disabled it wherever it is actually started —
docker/Dockerfile builds the bundle. phpunit.xml carries the switch now,
next to PULSE, TELESCOPE and NIGHTWATCH, and CLAUDE.md records why.

* refactor: one Meta connect flow instead of two kept in step by hand

The Facebook and Instagram-via-Facebook callbacks ran the same twenty-five
lines: the profile touch Meta's review wants, the granted-scope read and
the publish-scope refusal, the page walk, and the answer for a walk with
nothing to offer. They only matched because both were edited side by side,
every round, which is a guarantee nobody should be making by hand.

graphApi() moves to SocialController and reads the host by platform value,
so it serves every network rather than the two that had copied it, and
graphVersion() derives from it instead of reading config a second time.

select() stays as it is. The two differ in the middle — different identity
keys, different connect shapes — and folding them would be abstraction for
its own sake.

* fix: default Inertia SSR off, where this project already stands

Nothing in the repo starts an SSR process, so the shipped default was
describing a setup that does not exist. With it off the test env needs no
override of its own, and CLAUDE.md records that turning it on means
starting the process, not just flipping the env.

* fix: one rule for a refused portfolio, and a clock on the walk

Last round I made a refused /me/businesses mark the walk incomplete, on
the argument that refusing the index means "we could not look". That was
wrong in the case that matters most: an app without Advanced Access for
business_management gets that refusal on every single connect, so every
login on such an install lost auto-connect and every login without Pages
was told to retry forever. Self-hosted in Live mode is exactly that.

The rule that holds everywhere: a Page this login cannot enumerate is a
Page it cannot get a token for, so it was never connectable, and the list
of connectable Pages is complete. Only an unknown — a throttle, a hiccup,
a budget or a ceiling — leaves the walk unable to vouch for itself. Index
and edge now answer the same way, which is also what makes the two
readable together.

The per-request budgets were each bounded while their sum was not: ten
pooled rounds plus twenty-five cursor requests can outlive nginx's
fastcgi_read_timeout of 120s. The walk now carries a deadline and returns
what it has.

touchProfile exists only because Meta's review wants the call. It had no
timeout and no guard, so a hung /me could stall the callback to the
gateway timeout or fail a connect outright, over a response nobody reads.

* docs: the walk's contract changed under its own docblock

It still said any failure marks the walk incomplete, which stopped being
true when a refusal became an answer. A docblock describing an invariant
the code no longer holds is how this branch got two of its bugs.

* fix: put the whole callback inside the budget it advertises

meta_page_walk_seconds bounded the portfolio half of the walk and nothing
else. /me/accounts could paginate a hundred pages at fifteen seconds
each, and the Instagram lookups pooled in rounds that were themselves
serial — a portfolio with three hundred linked Pages is fifteen rounds,
after the walk had already spent its own budget. Both honour the deadline
now. The lookups skip rather than drop: the Page still connects, only its
handle and avatar arrive empty. The first request is always made; the
budget bounds what comes after it.

A Graph body that is valid JSON but not an object — a proxy answering
"throttled" — reached GraphError::isTransient, whose parameter is ?array,
and under strict_types raised a TypeError. That is an Error, so it walked
past both callbacks' catch(\Exception) and 500'd the popup instead of
showing a message.

Refusing a login before any listing has happened no longer borrows the
wording for "we found your Pages but not the permission to post to them".

composer run dev no longer starts an SSR process for SSR that is off, and
CLAUDE.md no longer claims nothing in the repo starts one, which
composer.json contradicted.

* fix: say what actually gets a Page token, per Meta's own reference

The Page node reference is explicit: access_token is "only returned if
the User making the request has a role (other than Live Contributor) on
the Page". Being an admin of the portfolio that owns a Page lists it but
does not grant that role, so the walk can surface Pages this login will
never get a token for.

The popup told those users to reconnect and accept every permission,
which cannot produce a Page role and so could never work. It now names
the role as well.

I rejected this in review on the grounds that a portfolio Page had been
published to successfully in the wild. That proved a token comes back
when the login holds a role, not that one always does.

* fix: one budget for the callback, not one per phase of it

The walk and the Instagram lookups each opened a full
meta_page_walk_seconds, on top of the profile touch and the permission
read, so the callback's worst case was several times the single bounded
budget config/trypost.php advertises. They share one deadline now, taken
once and passed down. META_PAGE_WALK_SECONDS joins .env.example.

An Instagram account described past that deadline arrived with no handle
and no name, and a lone one was then persisted with display_name null —
a blank, unidentifiable card. It falls back to the Page's own name.

Two docblocks were describing behaviour the code does not have.
/me/accounts is the base every other Page is added to, so running out of
budget there aborts rather than degrades, and the class now says so
instead of promising a partial list. GrantedPermissions justified
treating an absent permission as unknown but said nothing about a failed
request, which lands in the same place for a different reason.

* fix: a Pages throttle on a user token was reading as a refusal

Meta's BUC rate-limit table lists code 32 for the Pages API when called
with a User token. GraphError did not carry it, because until this branch
nothing in the app called a Pages surface that way — the publishers use
Page tokens, where the same throttle arrives as 80001.

The portfolio walk does: /me/accounts, /me/businesses and both edges are
read with the user token straight out of OAuth. So an ordinary throttle
came back as code 32, was classified as a confirmed rejection, and the
walk concluded this login simply reaches no portfolio Pages — vouching
for a list missing all of them and auto-connecting whatever /me/accounts
happened to hold. A rate limit was producing the exact silence the
complete flag exists to prevent.

* fix: a reconnect no longer loses its handle to a slow Graph

persistIdentity updates a reconnected card with whatever it is handed,
so a described-with-nulls card overwrote a working account's username
and avatar. Skipping the Instagram lookup — which the shared budget now
does whenever the walk spent it — produced exactly that card. A lookup
that never ran says nothing about a handle the account already has, so
those two keys are left out when it did not.

Refusing the portfolio index goes back to marking the walk incomplete. I
had it that way, reversed it, and this settles it: Meta's Page reference
returns access_token for a Page the login holds a role on, and a Page can
carry that token on a portfolio edge while /me/accounts omits it — which
is this branch's entire premise. So refusing one edge does say those
Pages are unreachable, but refusing the index says no edge was read at
all, and the Pages behind it may well have been connectable. Silently
vouching for a list without them is the original bug.

The budget also starts before the walk and now shapes each request's own
timeout, so no single call can outlive it by fifteen seconds.

composer dev:ssr was a slower alias of composer dev once the SSR process
came out of it.

---------

Co-authored-by: StoriaJames <james@storia.tech>
2026-08-26 10:42:59 -03:00
Hafiz Muhammad Moaz
91c3d86d86
Allow multiple social accounts per network via env (#286)
* feat: expose self-hosted mode to the accounts UI

SocialAccountObserver already bypasses the one-account-per-network
guard when trypost.self_hosted is true, but the frontend had no way
to know that and always collapsed a network to a single card once
any account existed - so self-hosted deployments could not surface
a second LinkedIn (or Instagram) connection even though the backend
would allow creating it.

* feat: allow connecting multiple accounts per network when self-hosted

NetworkConnectGrid always collapsed a network (LinkedIn profile/page,
Instagram standalone/Facebook) to a single card once any account
existed, with no way to trigger another OAuth flow - even though
SocialAccountObserver already allows unlimited accounts per network
in self-hosted mode. A self-hoster connecting their personal LinkedIn
profile had no path back to the connect flow to also add a company
page (or a second company page/showcase page).

Render one card per connected account instead of collapsing to the
first, and keep a standing "Connect another" card available for a
network's existing connections when self-hosted. Hosted mode is
unchanged: still one card per network, matching the backend's
still-enforced one-account-per-network limit there.

* test: cover the selfHosted prop on accounts and onboarding pages

Backend behavior for connecting a second identity per network in
self-hosted mode was already covered (LinkedInControllerTest,
NetworkUniquenessTest) - these just confirm the new prop the frontend
now depends on is actually present and reflects config correctly.

* style: apply prettier formatting

Pre-existing drift in this file unrelated to the selfHosted change.

* refactor: read selfHosted from shared Inertia props

The flag is already shared by HandleInertiaRequests, so the accounts
and onboarding controllers do not need to pass it again.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat: gate multiple social accounts with a dedicated env

Cloud cannot flip SELF_HOSTED, so one-per-network is now ALLOW_MULTIPLE_SOCIAL_ACCOUNTS (default false).

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: tighten multiple-account gates after review

Keep every connected identity visible, share occupiesNetwork, and return network_taken instead of a generic connect error.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: bind reconnect to the card and unique social identity

Reconnect now updates the selected account, and a unique index plus connectIdentity keep the same platform identity from being inserted twice.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor: build the OAuth URL before opening the popup

Keep the popup opener URL-only so reconnect query params are assembled at the call site.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor: drop dead social-account guards and slim the connect grid

Skip migration cleanup that production never needs, trust the platform enum in the observer, and move card theming out of the grid.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor: scope social reconnect to the current network

Drop dead instanceof/isset guards and filter reconnect targets in the query so a stale session cannot update another network's card.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor: slim connectable-identity filtering

Index OAuth identities by id so reconnect and occupancy use only/except instead of hand-rolled filters.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor: slim social identity persist helpers

Drop the unused occupiesNetwork exception and persist reconnects with update() instead of fill/save/fresh.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: keep reconnect updates on the original social card

Co-authored-by: Cursor <cursoragent@cursor.com>

* test: run the suite with multiple social accounts enabled

phpunit.xml forced ALLOW_MULTIPLE_SOCIAL_ACCOUNTS=false, overriding the
true value in .env.ci. That broke eight tests across Automation, MCP,
PostApi, RefreshExpiringTokens and VerifyUpcomingConnections which only
needed two accounts of one network as a fixture, not as a rule under test.

Match .env.ci instead. Every test that exercises the one-per-network rule
already sets the config itself; the accounts index test was the only one
leaning on the implicit default, so it now pins it.

* fix: align the multi-account fallback with the self-hosted default

allow_multiple_social_accounts fell back to env('SELF_HOSTED', false)
while self_hosted itself defaults to env('SELF_HOSTED', true). A
self-hosted install that never wrote SELF_HOSTED to its .env resolved to
false and silently lost multiple accounts per network on upgrade, which
is the opposite of what the documented fallback promises.

* fix: collapse duplicate identities before adding the unique index

Installs predating the index can hold the same identity twice: the
network guard was bypassed for multi-account installs and Pinterest
always created a fresh row. Creating the index on that data aborts
migrate mid-deploy.

Keep the newest row per identity and move its post_platforms over before
dropping the duplicates - the FK is nullOnDelete, so deleting outright
would orphan drafts and scheduled posts.

* fix: refuse a reconnect that authorized a different identity

connectIdentity overwrote platform_user_id with whatever the provider
returned, so reconnecting a card while signed into another account
repointed the row - and every draft and scheduled post bound to it - at
a stranger. LinkedIn guarded this at the controller and Facebook via its
filtered page list; nothing covered X, TikTok, Threads, Discord,
Bluesky, Mastodon, Pinterest, Instagram or Telegram.

Enforce the identity match at the single choke point every connect flow
goes through. Every call site already maps NetworkAlreadyConnectedException
to network_taken, so the refusal surfaces without new plumbing.

Also restore the null-platform guard in the observer: occupiesNetwork
type-hints a non-nullable Platform, so a row without one died with a
TypeError instead of the database's NOT NULL error.

* fix: filter connectable identities on every picker step

YouTubeController::select re-fetched the channels and matched the posted
id straight off the raw list, unlike callback and selectChannel. With a
live youtube_oauth session it let a POST name any channel the Google
account owns and bind it to the reconnect target. It also read the
reconnect from the session while the connect below it read
youtube_oauth.reconnect_id, so the two could disagree - pass the
resolved account through instead.

filterConnectableIdentities also short-circuited in multi-account mode,
and the unique index is scoped to platform rather than network. That let
one Instagram account connect twice, once directly and once via
Facebook, publishing every Instagram post to it twice. The existing
except() already spans networkPlatformValues(), so dropping the
short-circuit closes it.

* refactor: type the connect cards and drop the dead accounts grid

The cards computed inferred account as a required ConnectedAccount and
then pushed undefined onto it (TS2345). CI only runs eslint so it stayed
green, but vue-tsc and editors flag it.

SocialAccountsGrid is referenced nowhere; its reconnect button was
updated in this branch without passing the card id, which would have
been a bug had anything rendered it.

* test: keep the suite on the cloud one-account-per-network default

CI runs the Cloud build, so the suite baseline should be the Cloud
default rather than the self-hosted one. Put phpunit.xml and .env.ci
back to false and make the eight tests that merely need two accounts of
one network as a fixture opt in for themselves.

This also un-deads the config()->set(true) calls the branch had already
added to AuthenticationTest, SyncAccountUsageTest, HasUsageTraitTest and
SocialAccountObserverTest, which the forced true had turned into no-ops.

* fix: connect standalone instagram instead of reopening the picker

The picker emits an already-resolved connect method, but this branch
rewired @select from openOAuthPopup to startConnect. startConnect sends
a bare 'instagram' straight back into its own picker branch, so choosing
"Instagram" closed the dialog and immediately reopened it - the OAuth
window never opened and the standalone flow was unreachable. Only the
via-Facebook button still worked.

Split the URL-opening tail out of startConnect and let the dialog call
that directly.

* fix: reject a telegram reconnect before burning the connect code

The nonce was consumed before connectIdentity ran, so posting /connect in
the wrong chat spent the one-off code and forced the user to generate a
new one. Check the identity first and report wrong_chat instead of
network_taken, which told them to disconnect an account when the real fix
was posting in the channel they were reconnecting.

* fix: leave one target per post when merging duplicate accounts

post_platforms has no unique on (post_id, social_account_id), so a post
holding a row per duplicate account ended up with two enabled rows aimed
at the surviving account and would publish to it twice. Keep one row per
post, preferring a published one so history survives.

* refactor: collapse the repeated connect-flow boilerplate

Four shapes were copy-pasted across the connect controllers:

- the session + permission guard opening 16 actions, now connectWorkspace()
  throwing a ConnectPopupException that renders the popup itself
- the reconnected/connected ternary in 13 places, now connectedCallback()
- the "nothing left to connect" branch in 4 places, now
  noConnectableIdentities()
- validatedReconnectId() re-querying what reconnectAccount() already does

Facebook, Instagram-via-Facebook and YouTube also re-queried the reconnect
account three or four times per callback; it is resolved once and passed
down. The three GET pickers skipped the manageAccounts check their POST
siblings had, and pick it up from the shared guard.

Drops the color key from connectableOptions and the matching frontend
field - nothing read it. Platform::color() stays; the disconnection
emails use it.

* fix: keep an expired connect popup out of the error log

ConnectPopupException escapes to the framework handler so it can render
itself, which also meant report() ran first: every session_expired and
workspace_not_found popup filed an ERROR and a Nightwatch issue for what
used to be a silent return. A stale popup is a normal outcome, so it now
implements ShouldntReport.

The Mastodon and Threads guards also cleared their provider session
after connectWorkspace(), so a workspace that vanished mid-flow left the
client secret and the OAuth state behind. Clear first, then resolve.
clearMastodonSession() no longer touches social_connect_workspace -
whatever closes the popup already does.

* fix: stop telling users to disconnect an account that is not the problem

Two flows reused popup_callback.network_taken - "This workspace already
has an account for this network. Disconnect it first." - for situations
where that is neither true nor actionable:

- reconnecting a card while signed into a different account on the
  provider, now wrong_account
- an empty picker in multi-account mode, where every page or channel on
  that login is simply already connected, now all_connected

NetworkAlreadyConnectedException carries the message key so the catch
sites stay one line. handleCallback() also drops its $platform argument;
it read $this->platform for the reconnect lookup and the identity filter
either way, so a caller passing a different platform would have scoped
the lookup to the wrong network.

* refactor: filter linkedin identities with the shared helper

The picker hand-rolled its own reconnect narrowing because the profile
and the pages arrive in two different shapes. Flatten them into one pool
of LinkedIn identities, run the shared filter, and split them again for
the view - the same path Facebook, YouTube and Instagram already take.

Side effect worth having: the picker previously only narrowed on a
reconnect, so it would offer an identity that is already connected and
only fail once the user picked it. It now hides taken identities up
front and says so when nothing is left.

* fix: keep the linkedin picker's own empty state

Routing the picker through the shared filter made every empty pool look
like "nothing left to take", including the pool LinkedIn never filled.
A self-hoster running pages-only who administers no page was told the
network was already connected, or that every account on the login was
taken - both false - and the picker's own "you are not an admin of any
LinkedIn page" state became unreachable.

Only treat it as taken when filtering is what emptied it. Splitting the
pool back also compared the person id loosely on one side and strictly
on the other; one predicate now drives both.

Threads had two forget() calls for a key the top of the action already
clears, and YouTube's picker resolved the reconnect account twice on the
failure path.

* fix: keep the enabled row when collapsing duplicate post targets

SyncPostPlatforms seeds a disabled post_platforms row for every account
in the workspace, so the usual duplicate is one row the user actually
checked next to one they never saw - both pending, both created in the
same second. Ordering only by published-then-newest made that a coin
flip, and PublishPost iterates enabled() only, so half the time a
scheduled post would silently stop reaching that account and take its
caption and per-platform meta with it. This runs once against production
data and the dropped row is gone, so enabled now beats disabled.

Also: the empty-pool exit from the LinkedIn picker was the only one
leaving linkedin_pending - and its tokens - in the session. The
rationale comments move to the docblocks they belong in, and usePage()
comes out of the cards computed.

* fix: stop the migration destroying publish history and automations

Two ways the one-shot merge lost data that cannot be rebuilt:

Surplus published post_platforms rows were deleted. Two duplicate
accounts really could each have published, and each row carries the
platform_post_id for a live post on the network - dropping one leaves
that post unmanageable and invisible to metrics. The docblock claimed
published beat everything; now the code does, and only unpublished
repeats collapse.

Automation nodes persist social_account_id inside a JSON column with no
foreign key, so deleting the loser left RunGenerateNode skipping that
target, or generating nothing at all when it was the node's only
account. The ids are rewritten - current and legacy shapes both - and
entries the merge just turned into duplicates are collapsed.

Ordering is now total (null created_at sorts oldest on every engine,
then id) so a rehearsal on a replica keeps the same rows as the real
run. The LinkedIn picker also passes onboardingProgress inline: it
clears linkedin_pending on the empty path, and a deferred reload would
re-GET the route and swap the empty state for a session-expired popup.

* fix: make the identity merge auditable and stop a second delivery

Self-hosted installs run this unattended and it cannot be undone, so
each collapsed group now logs the workspace, the identity, which row was
kept, which were dropped, and how many post_platforms and automations it
touched. down() says plainly that it drops the index only.

Two narrower fixes:

A post holding a published row plus an enabled unpublished row for the
same account kept both, and PostPlatform::scopeEnabled() filters on
`enabled` alone with no status check - so a republish would deliver the
same content to that identity twice. Once a published row exists, every
unpublished repeat goes.

The automation dedupe ran on every automation in the workspace, not just
the ones the merge rewrote. A node legitimately holding two entries for
one account under different content types would be collapsed to
whichever came first in the array. It now runs only where an id was
actually substituted.

* test: rehearse the identity merge against a messy database

Every test on this migration so far covered a case someone thought to
write, which is why three separate review rounds each found a defect the
earlier ones missed. This builds a deliberately messy database instead -
three workspaces, four networks, one to three copies of each identity,
posts mixing published, pending and failed rows across the duplicates
with enabled flags varying, and automations referencing them in both the
current and legacy JSON shapes - then runs the real migration and
asserts what must be true afterwards rather than what happens to a
particular fixture.

Invariants: no duplicate identity survives, no published row is ever
destroyed, no post ends up enabled twice against one account, nothing in
post_platforms or automations points at a deleted account, and the
newest row of each identity is the one kept.

The generator is seeded, so a failure reproduces, and it asserts its own
output is adversarial - roughly nine duplicate groups and fourteen
published rows - so it cannot quietly degrade into passing on an empty
problem. Verified by mutation: dropping the automation repoint, the
published guard, or the repeated-target collapse each fails exactly the
invariant that covers it.

* fix: stop the youtube picker refetching itself into a cleared session

HandleInertiaRequests defers onboardingProgress for anyone mid-onboarding
- exactly the people connecting their first accounts - so Inertia
re-GETs the picker route right after it mounts. For Facebook and
Instagram that re-entry is harmless and deliberately left deferred, but
YouTube calls the Google API again, and fetchChannels() turns any
failure into an empty list that clears the connect session and swaps the
mounted picker for an error the user cannot retry from. Same guard the
LinkedIn picker already got.

LinkedIn also answered a reconnect that authorized a different identity
with "Page not found", including in the person branch where no page is
involved. Every other platform says wrong_account, which this PR added.

* refactor: drop the unreachable youtube channel picker

Google's own delegation screen already lists every channel on the
account and makes the user pick one before it issues the token, so
channels?mine=true always answers with that single channel and
count($channels) === 1 always won. The picker behind it was never
reached - its Vue page was deleted back in 7c00c338 (January) and
nothing broke, which is the clearest evidence it was dead.

Removes selectChannel(), select(), both routes, the youtube_oauth
session payload and the tests that drove them. If Google ever does
return more than one, the callback connects the first and logs a warning
rather than routing to a screen that no longer exists.

* fix: serialize connects so two popups cannot seat one network twice

The observer's occupiesNetwork() is a check-then-insert with nothing
holding the gap, and the new unique index covers the identity, not the
network. Two tabs finishing OAuth at the same moment for *different*
identities on one network both passed the exists() check and both
inserted, leaving a Cloud workspace with the two accounts the rule
exists to prevent. The same-identity race was already safe - the unique
violation is caught and re-queried.

A database constraint cannot hold this: allow_multiple_social_accounts
is a runtime flag, so the rule is on for Cloud and off for self-hosted,
and an index cannot read config. Lock per workspace and network instead,
the way markAsDisconnected() and ConnectionVerifier already do.

This covers connectIdentity(), which every OAuth flow and the Telegram
action go through. A direct create() still answers to the observer
alone, and a self-hosted install running file cache across several nodes
locks per node.

* fix: handle a busy connect lock on the telegram path

Every other caller funnels LockTimeoutException into its generic
\Exception catch and closes the popup with error_connecting. Telegram
has no such catch, so the new lock could 500 the webhook - and because
the nonce is spent before connectIdentity runs, Telegram's retry of the
same update short-circuits on the consumed code and returns without
dispatching anything. The dialog would spin forever on a code that can
no longer be used.

Also restores coverage the picker removal dropped: the deleted select
tests were the only ones driving a multi-channel response, so nothing
exercised the reconnect narrowing to its own card, or multi-account mode
skipping an already-connected channel. Both are back against the
callback, and removing the narrowing in filterConnectableIdentities
fails them.

* fix: stop the instagram login seating an account already held via facebook

filterConnectableIdentities() drops every identity already connected on the
network, which is what keeps one Instagram account from being seated twice
under its two platforms. Every flow that persists an identity ran it except
the direct Instagram Login callback, so the guard only held in one direction:
InstagramFacebookController refused an account already connected as
`instagram`, but the reverse was allowed through.

With multiple accounts per network enabled the observer's network check is
bypassed and the unique index does not span platforms, so authorizing the
same account through the direct flow created a second row. Both then seed a
post_platform row and the post goes out twice to one account.

* fix: name the real reason when a linkedin profile reconnect switches member

Reconnecting a card narrows the authorized identities to that card's own, so
authorizing a different LinkedIn login empties the pool. selectIdentity()
reported that as "Page not found." for every card, including personal
profiles where no page was ever involved.

A profile reconnect has no page to be missing: an empty pool there can only
mean this login is a different member. Say so with the wrong_account wording
select() already uses for the same condition. Page reconnects keep
page_not_found, where the organization really can be absent from the login.

* fix: surface the busy telegram connect instead of a generic failure

The connect lock timing out dispatches its own 'busy' reason so the dialog
can tell the user to retry, but the dialog only mapped network_taken and
wrong_chat and fell back to error_generic for everything else. The reason
reached the browser and died there, leaving "Could not start the connection"
for a case that just needs another moment.

* test: cover reconnect on every flow that gained it

rememberConnectSession() gave Instagram, TikTok, Threads, Mastodon and
Bluesky a reconnect path they did not have before — TikTok had been actively
clearing social_reconnect_id on connect — and none of them had a test for it.
Facebook, LinkedIn, YouTube, X, Discord, Pinterest and Telegram already did.

Each now covers both halves: authorizing the same identity refreshes the
existing card and reports it as a reconnect, and authorizing a different one
is refused with wrong_account instead of quietly seating a stranger on the
card and every post scheduled against it.

* fix: repair what a reconnect leaves behind when it cannot proceed cleanly

Two things connectIdentity got wrong once the reconnect path existed.

A reconnect through the other variant of a network moves the card to the new
platform — same identity, different API flavor. Post targets carry their own
platform snapshot, and that snapshot picks the publisher, the queue and the
scopes checked before publishing. Left behind, it failed every pending post on
a permission the account no longer needs: an Instagram card moved to the
Facebook variant still demanded instagram_business_content_publish and stopped
with "Missing permissions". Pending targets now follow the card and reset a
content type the new platform cannot publish; published targets keep theirs,
since they record what really went out under a platform_post_id from that API.

The network lock timing out also arrived as a raw LockTimeoutException, which
every OAuth callback filed through its generic catch: an error log and "Error
connecting account" for the exact race the lock exists to absorb. It now
carries a busy messageKey through the branch each flow already handles, the
same way the Telegram path already reported it.

* refactor: resolve the linkedin reconnect card once per select

select() already looked the card up before deciding whether the chosen
identity matches it, then connectPerson() and connectOrganization() looked it
up again on their own — two identical queries per submit, and two places that
could disagree about what is being reconnected. The caller passes what it
already holds.

* test: render the grid's multi-account branch

phpunit.xml forces ALLOW_MULTIPLE_SOCIAL_ACCOUNTS false and no browser test
overrode it, so the card the flag exists to add never rendered anywhere. The
pair pins both sides: a taken network offers no second card when multiples are
off, and offers one when they are on.

* test: pin why the linkedin select guards exist

connectIdentity() already refuses a mismatched reconnect and answers with the
same wrong_account message, so every existing test passes with the two guards
in select() deleted — which is exactly how they would get deleted. What they
actually buy is skipping the avatar download that building the connect payload
runs first.

Both now assert the fetch never happens, so the guards fail loudly instead of
looking redundant.

* fix: carry retrying targets through a variant move, atomically

Two holes in the move added a commit ago.

It only carried pending targets, but a retrying one is not finished either —
the publish job reschedules itself and reads the snapshot fresh on the next
attempt, so leaving it behind meant it retried against the old variant until
it exhausted its budget on a permission the account no longer needs. Failed
and published targets stay put; a publishing one has a job mid-flight already
working from the snapshot it read.

The card and its targets also moved in three separate statements, so a crash
between them left exactly the split this was meant to close. They share a
transaction now.

* chore: drop the dusk selectors nothing reads

Laravel Dusk is not installed — no laravel/dusk requirement, no DuskTestCase,
no browse(). Browser tests run on pest-plugin-browser driving Playwright, and
its @selector resolves to data-testid. The 45 dusk attributes left across 18
components selected nothing.

CLAUDE.md was the reason they kept coming back: it told every agent to add
them. Its browser-testing section now describes the setup that exists —
data-testid targeting, the wait helper these tests need because assertions do
not auto-wait on SPA paint, and why BrowserTestCase keeps Vite real.

Verified before removing: every @selector used in tests/Browser resolves to a
data-testid, seven of them through bound :data-testid, so none depended on a
dusk attribute.

* chore: drop the last one-account-per-network helper

hasConnectedPlatform() has no callers left anywhere — app, tests, views or
routes. It sat directly above getSocialAccount(), which this branch already
removed, and is the same leftover from when a workspace could hold one account
per platform.

---------

Co-authored-by: Paulo Castellano <paulo@castellanos.llc>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-25 07:28:14 -03:00
Paulo Castellano
96ec995fcd
Remove the X API reads that buy nothing (#299)
* Stop paying X for token checks the refresh already proves

RefreshSocialToken verified rotating-refresh platforms instead of
refreshing them. On a still-valid token verify() only called the platform's
verify endpoint and left token_expires_at untouched, so the account stayed
inside RefreshExpiringTokens' 30-minute window and was re-read every 15
minutes until the token actually died.

On X that endpoint is GET /2/users/me, billed as a "User: Read" ($0.010 per
resource under X's pay-per-usage pricing). Simulating 24h of the scheduler
against one connected X account: 36 billed reads per day, 24 of which
renewed nothing, plus 180 minutes per day sitting on an expired token
between expiry and the next tick.

Refresh the token outright instead. A provider that hands back a fresh token
has already confirmed the credential — it rejects a revoked one with a 4xx,
which TokenRefreshClient maps to TokenExpiredException — so the verify call
adds cost and nothing else. Record the confirmation in last_verified_at, and
let the daily sweep trust it for 12 hours the way VerifyUpcomingPostConnections
already does, so it stops re-reading accounts a refresh just proved valid.

Same simulation after the change: 0 billed reads, 0 minutes expired.

Two existing tests asserted the old policy (verify-first, refresh_token left
unrotated) and now assert the new one. The rotation test still guards what
made that policy attractive: a proactive rotation must not trip a
false-positive disconnect.

* Don't disconnect an account whose access token still works

Refreshing instead of verifying removed a safety net the old verify-first
path had: if the refresh is rejected, the account was marked TokenExpired
outright. But a rejected refresh does not mean the connection is dead.

X and LinkedIn single-use their refresh_token, so a token a concurrent
refresh already consumed comes back 4xx while the current access_token keeps
working. An account with no refresh_token at all fails even earlier, without
a single call being made. Verified against main: both cases stayed Connected
before, and became TokenExpired after. PublishToSocialPlatform hard-fails
every post for a TokenExpired account, so this killed posts the access_token
would have published, up to 30 minutes before the token was actually due to
expire — and emailed the owner a disconnect notice for it.

Fall back to verifying the access token before disconnecting. This is the
only path in the job that reaches the billed verify endpoint, and only after
a refresh has already been rejected, so the healthy path stays at zero reads
(re-confirmed: 0 reads and 0 expired minutes across a simulated 24h). A
failure that can't be attributed to the token — platform down, network blip —
leaves the account alone instead of disconnecting it on noise.

Also covers two gaps found while reviewing: a lock-skipped refresh must not
record a verification it never performed, and a platform with nothing to
refresh must not be recorded as verified either. Both already behaved
correctly; they now have tests so the daily sweep can't start trusting a
stamp nobody earned.

* Only record a verification something actually proved

Review of the previous two commits turned up four issues, all in code they
introduced.

The stamp lived inside ConnectionVerifier::refreshToken(), which
refreshThenVerify() also calls — there the refresh can succeed and the verify
that follows still fail. The stamp was already written by then, vouching for a
credential nothing confirmed. Normally harmless, because the caller marks the
account TokenExpired and recentlyProvenValid() only skips Connected ones, but
markAsTokenExpired() silently no-ops when its status lock is held by a
concurrent publish. The account then stays Connected with a fresh stamp, and
both skip-windows wave it through: 40 minutes before publishing, 12 hours in
the daily sweep. Move the stamp to the caller that owns the outcome.

TokenRefreshClient classifies on HTTP status alone and never inspects the
body, so a 200 carrying an empty access_token is stored as-is and was then
recorded as healthy. (A missing key rather than an empty one can't get that
far — access_token is NOT NULL, so the write throws first.) Guard on a filled
token.

The fallback added in 9775578 called verify(), which for an already-expired
account — RefreshExpiringTokens selects those too — runs refreshThenVerify()
and re-sends the refresh_token the provider just rejected, and on Bluesky
re-runs the password re-auth AT Proto rate-limits per account. Its own
docblock claimed it only reached the verify endpoint. Add
verifyAccessToken(), which checks the stored token and nothing else.

VerifyWorkspaceConnections read last_verified_at without ever writing it, so
an account it had just confirmed healthy still burned a fresh call minutes
later when a post entered the risk window. Stamp on its success path too.

Also corrects the RefreshExpiringTokens docblock, which still described the
verify-first behaviour removed in c0d0d57.

Re-ran the 24h whole-scheduler simulation: still 0 billed reads, 0 expired
minutes, account Connected at the end.

* Fall back to the token a concurrent refresh persisted

The fallback added in 9775578 judged the in-memory access_token, which is
exactly the one that is stale when a refresh loses a race.

X single-uses the refresh_token, so when two refreshes overlap the loser's
token comes back 400 invalid_grant. Its in-memory instance still holds the
pair the winner has already rotated away, so verifying it 401s and the
account is marked TokenExpired — while the row in the database holds a
perfectly healthy token the winner just wrote.

Verified against main: a concurrent rotation leaves the account Connected
there and TokenExpired here. refreshThenVerify() already handles this by
reloading and retrying with whatever was persisted; the new path skipped that
because it never went through refreshThenVerify.

Reload before judging. This is the failure path only, so the healthy path is
untouched — the whole-scheduler simulation still reports 0 billed reads, 0
expired minutes, Connected at the end.

* Don't record a verification when the lock skipped the refresh

refreshToken() returns normally — no exception — when another process already
holds the per-account lock, so the caller can't tell "refreshed" from "did
nothing". Moving the stamp out of the verifier in 4d2795a lost that
distinction: RefreshSocialToken stamped last_verified_at on a run that made
zero HTTP calls.

The account is then vouched for by nobody: the daily sweep skips it for 12
hours and the pre-publish check for 40 minutes. If the concurrent refresh
also failed, nothing ever confirmed the credential.

Have refreshToken() report whether it actually ran. Callers that ignore the
return value are unaffected.

The test meant to cover this called refreshToken() directly rather than going
through the job, so it kept passing while the job path was broken — it now
exercises the job and asserts no HTTP call was made.

* Close the two concurrency gaps the refresh path leaves open

Both are pre-existing, but this branch raises the exposure to them from 12 to
16 refreshes per account per day.

The per-account lock lasted 30 seconds, exactly the HTTP client's default read
timeout — so a refresh could outlive the lock that protects it. Bluesky is the
worst case, refreshing with two sequential calls (refreshSession, then the
createSession re-auth), each bounded by connect + read timeouts: up to ~80
seconds under one 30-second lock. Once it lapses, a second process refreshes
with the same single-use refresh_token and one of the two is rejected. Name
the TTL, set it past the ceiling, and write down the invariant so a future
slower refresh doesn't quietly break it.

RefreshSocialToken was not unique. RefreshExpiringTokens re-selects an account
until token_expires_at moves, and that only happens once the job runs — so a
queue more than one tick behind stacked a job per tick for the same account,
each rotating a single-use refresh_token again for nothing and widening the
gap where a worker death loses the pair. Key it by account like
VerifyUpcomingPostConnections already does.

Cadence is unchanged: the whole-scheduler simulation still reports 16
refreshes, 0 billed reads and 0 expired minutes over 24h.

* Correct which providers actually single-use their refresh_token

Checked each provider's official documentation rather than carrying the
assumption forward.

LinkedIn does not rotate. Its refresh docs are explicit: "the lifespan or Time
To Live (TTL) of the refresh token remains the same as specified in the
initial OAuth flow (365 days)" — the same token comes back with a decreasing
refresh_token_expires_in, and only the access token is reissued. The claim
that it single-uses the token predates this branch, but a docblock added here
repeated it.

Bluesky does rotate, and belongs in the list instead: com.atproto.server
.refreshSession declares refreshJwt as a required output field, so every
refresh mints a new one.

Verified alongside, all matching what the code already does:

  X          access token 2h, refresh single-use with rotation
  Bluesky    refreshJwt rotates; createSession is rate-limited per handle
             (30/5min, 300/day), which the fallback re-auth path shares
  TikTok     access 24h, refresh 365d, "may be different — you must use the
             newly-returned token", which refreshTikTokToken does
  LinkedIn   access 60d, refresh 365d fixed, not rotated
  Google     does not rotate on refresh; 100 refresh tokens per account per
             client, so the higher refresh rate on YouTube carries no
             rotation risk

* Read X post metrics from the timeline that already returned them

Analytics fetched the account's timeline for post ids, then turned around and
looked the same ids up again through GET /2/tweets purely to read the
public_metrics the first request could have returned. The timeline call asked
for start_time, end_time and max_results — never tweet.fields.

Both endpoints bill per Post returned, so the second pass claimed the same
resources a second time, took a second round-trip, and spent a second slice of
the same rate limit. For an account with 250 posts in range that is 6 requests
where 3 will do.

The saving is in round-trips and rate limit rather than dollars: X deduplicates
a resource within a 24-hour UTC window, so the second read of an id already
read that day is not charged again. But the docs call that a soft guarantee
that "may result in resources not being deduplicated" — this stops leaning on
it for 250 resources per analytics load.

Behaviour is unchanged: same totals, same 5-page ceiling, same empty result
when the account posted nothing in range. The page cap is now a named constant,
since it bounds what one load can cost as much as how long it takes.

Adds the first tests for XAnalytics::getMetrics, covering the totals, the
pagination, and that the metrics arrive on the timeline request.

* Cover the analytics paths a happy-path test walks straight past

Three gaps in what the previous commit's tests actually assert:

A timeline page failing mid-pagination breaks out of the loop and returns
whatever was collected. Nothing pinned that — partial data beats an exception
on a dashboard someone is looking at, and a future refactor could quietly turn
it into one.

A post can come back without public_metrics. The accumulator defaults each
metric to 0, so it contributes nothing instead of erroring, which also wasn't
covered.

An account with no posts in range returns [] rather than a list of zeros, so
the UI can tell "nothing posted" from "posted, no engagement".

Also makes the routing test's mock return explicit. Without andReturn, Mockery
hands back a falsy default for the new bool return type, so the assertion about
routing was passing while silently exercising the lock-skipped branch. The test
still asserts only what its name claims, but no longer depends on a mock
default to get there.

* Close five issues an independent review found in this branch

All five sit in code these commits introduced.

Instagram and Threads must fail loudly. Their long-lived token is extended in
place and cannot be renewed once it expires, and RefreshExpiringTokens picks
them up a full day ahead precisely so there is time to react. The fallback
added in 9775578 applied to them too, so a permanently rejected extension on a
token that still reads left the account Connected — the daily sweep passed as
well, since verify() succeeds on it — and the owner learned about it only after
the token died unrecoverably, while every tick retried the rejection for 24
hours. The fallback now applies only to platforms that rotate a refresh_token,
which is what it was written for.

The lock went the wrong way. Lengthening it to 120s in 5001c9a treated the
scheduler as the only caller, but publishers wait on the same lock: one left
behind by a worker that died mid-refresh makes refreshToken() return false, and
the publisher falls through and publishes with an expired token. That window
was 30s and had become 120s. Bound the calls instead — a token endpoint answers
in milliseconds, and 8s read / 4s connect keeps even Bluesky's two sequential
calls under a 30-second lock, back to where main had it.

Reloading the account can throw. $this->account->refresh() sat outside the try
in the fallback, and an exception raised inside a catch block is not caught by
a sibling catch. With tries = 1, an account deleted mid-run put the job in
failed_jobs. VerifyUpcomingPostConnections guards this same race explicitly.

An empty access token was detected but not acted on. recordVerification()
declined to stamp it, yet the refresh still counted as a success — and the
refresh method had already pushed token_expires_at two hours out, so the
account left the window looking healthy while every publish 401d. Mark it
expired, which is what verify() used to do on the same input.

refreshToken() claimed "whether a refresh actually ran" but returned true for
platforms whose match arm does nothing. Only recordVerification() re-checking
hasTokenRefreshFlow() kept that from mattering. The guard is now explicit and
the contract true at the source.

Also settles the tweet.fields question against the live API rather than the
docs, which contradict each other: the OpenAPI spec names the parameter
post.fields, while the Fields guide and every example use tweet.fields. Both
are accepted and both return public_metrics. An unrecognised name returns 200
and silently omits the field — no error — so the name being right is load
bearing, and it is.

* Guard the match that no longer has a default arm

Dropping `default => null` from refreshToken() in af379af made the return value
honest, but it also turned a missing case into an UnhandledMatchError at
runtime. The arms and hasTokenRefreshFlow() currently agree on the same nine
platforms, and nothing enforces that: adding a platform to the predicate
without an arm would fail in production, on a queue worker, for one platform's
accounts only.

The test walks every platform claiming a refresh flow and fails by name if the
match has no arm for it. Verified it catches the real thing by temporarily
adding Facebook to the predicate — it failed with "facebook claims a refresh
flow but refreshToken() has no arm for it" — rather than trusting a green run
on code that already agrees with itself.

* Make the per-platform guard fail on a broken client chain

The test added in the previous commit swallowed every Throwable except
UnhandledMatchError, so it proved a match arm existed and nothing more. Routing
all nine refresh methods through refreshHttp() in af379af rewrote how each one
builds its request, and this test would have passed just the same if one of
those chains no longer worked.

It now fails on any exception, naming the platform, and asserts each refresh
actually put a request on the wire — a chain that breaks during construction
raises before anything is sent, so an empty recording is the signal.

Verified by breaking Pinterest's chain on purpose: "pinterest refresh threw
BadMethodCallException: Method PendingRequest::withHeadersTypo does not exist."
All nine send their request with the chains as they stand.

* Stop trading a recoverable failure for an unrecoverable one

A max-effort review found six issues, three of which undo a trade the previous
round got backwards.

Every refresh method wrote the response straight over the stored credential:
`'access_token' => data_get($data, 'access_token')`. A 200 carrying no token
therefore destroyed a working one — and on Instagram and Threads, where
refresh_token is set to the same value, both halves at once. The blank() check
added in af379af only noticed after the damage was persisted, then marked the
account TokenExpired, which RefreshExpiringTokens no longer selects — so one
glitchy-but-successful response emailed the owner and forced a manual
reconnect. Guard before the write instead, treat it as the platform
misbehaving, and the stored pair survives for the next tick to retry. The
detection branch downstream is now unreachable and gone.

Tightening the refresh timeout to 8s was the wrong fix for the lock problem.
refreshHttp() is shared with 24 publish and analytics call sites, and for X,
Bluesky and TikTok the refresh_token is single-use: abandoning a request the
provider has already processed loses the rotated pair permanently and costs the
user a reconnect. Giving up sooner makes that more likely, not less. Restore
generous timeouts and put the lock back above them. The cost of erring long is
that a worker dying mid-refresh holds the lock while a publish falls through
and retries — recoverable, unlike a lost rotation. Both constants now say which
way they are wrong on purpose.

The hasTokenRefreshFlow() guard in recordVerification() was dead: refreshToken()
already returns false for those platforms, so the branch was never entered. The
test claiming to cover it calls refreshToken() directly and never reaches it.

The command reported "Dispatched N" for a number it cannot know. dispatch()
returns a PendingDispatch whether or not ShouldBeUnique discarded it, so the
count overstated itself during exactly the backlog someone reads that line to
diagnose. It now reports accounts in the window, which is what it actually
measured.

Not changed: the daily sweep still skips accounts a refresh keeps fresh. That
is the deliberate decision this PR is built on — a refresh replaces the access
token rather than inspecting it, so there is nothing left for a billed read to
confirm.

Re-verified live after the changes: the real job still rotates the token
against api.x.com and leaves the account connected.

* Cover the tokenless-200 guard on every platform, not just X

The guard added in the previous commit protects nine refresh methods; only X
had a test. Each provider reads a different field name out of the response, so
a regression would land on one platform at a time and the suite would stay
green for the other eight.

The test drives every platform claiming a refresh flow through a 200 that
carries no token, and requires each to refuse with PlatformUnavailableException
— nothing is provably dead, so the next tick should retry rather than anyone
being disconnected — while leaving the stored credential untouched.

Verified it fails usefully by dropping the guard from one platform: 'threads
should refuse a tokenless 200 cleanly, got QueryException: null value in column
access_token violates not-null constraint'. Without naming the platform the
failure reads as an unrelated database error, since the write also poisons the
surrounding transaction.

* Stop the fallback from reading every failure as good news

A fifth review found the concurrent-refresh test passing without ever
exercising what it claims to cover, and the reason it could is a real bug.

access_token is an encrypted cast. The test wrote the winner's pair with
DB::table()->update(), which stores plaintext, so reading it back raised
DecryptException — and accessTokenStillWorks() caught Throwable and returned
true. Green test, zero coverage of the recovery that justifies the method
existing. It now writes through the model and asserts the verify call actually
happened; removing the reload makes it fail.

The catch is the bug. Treating any non-TokenExpiredException as "the token is
healthy" means an APP_KEY rotation, a corrupted column, or an UnhandledMatchError
from a newly added platform leaves the account Connected forever while every
publish hard-fails, and nobody is told. It now names the outcomes that earn the
benefit of the doubt — platform down, network dropped, account deleted mid-run
— and lets the rest surface. Loud is right here: an APP_KEY rotation breaks
every account at once, so failing the job where an operator sees it beats
disconnecting every user.

Refusing to persist a tokenless 200 also had no way out. The account kept
retrying every 15 minutes forever, and the daily sweep counts
PlatformUnavailableException as verified, so it was never disconnected and
never reported. Now the retry only continues while there is a live token
behind it: once that expires and renewal still fails, the connection is dead
in practice and says so.

Also corrects the refreshHttp() docblock, which claimed a blast radius the
private method does not have — refreshToken() is what those 24 call sites
reach — and records in VerifyWorkspaceConnections that short-TTL platforms
never being re-verified is the intended consequence, not an oversight.

* Stop a bad hour at the provider from disconnecting anyone

The escalation added last round was wrong, and two neighbouring paths had the
same shape of bug.

Marking the account expired whenever a PlatformUnavailableException hit an
already-expired token looked like it closed a silent-rot gap. But that
exception is what TokenRefreshClient raises for 5xx, 429 and connection
timeouts — so X rate-limiting for forty minutes around a two-hour token's
expiry disconnected the account, emailed the owner, and hard-failed every
scheduled post, with only the daily sweep to undo it. The rot it was meant to
prevent surfaces at publish time anyway. Reverted: a transient failure never
disconnects.

refresh_token was left unguarded when access_token was hardened. data_get()
only falls back when a key is absent, so a provider answering with an explicit
"refresh_token": null wiped the stored one — and the next tick then threw "no
refresh token available" without making a single call. Guarded in the same four
places, falling back on blank rather than on missing.

A held lock reported "nothing refreshed" even when the token was already dead,
handing the caller a credential it knew was expired. The publisher posts with
it, takes a 401, and PublishToSocialPlatform finalises the post as failed and
disconnects the account — over a lock a dying worker left behind, for the two
minutes it survives. It now says transient, which is what a refresh someone
else is already running actually is.

VerifyWorkspaceConnections promoted TokenExpired accounts back to Connected on
verifyAccount()'s return value, which is also true for "could not check, don't
disconnect". An unreachable platform therefore told owners their reconnect had
worked when nothing was verified. Promotion moved next to the successful
verify.

Each of the four is pinned by a test, and each test was checked by reverting
the fix and confirming it fails.

* Keep a lock collision off the analytics page

Round seven found one real regression from round six, one consistency gap, and
one stale comment.

Reporting lock contention as transient was right for the publish path, which
reschedules, but analytics calls refreshToken() bare and AnalyticsController
has no try/catch — and there is no renderable handler for
PlatformUnavailableException. A user opening analytics for an account whose
token expired while the scheduled job held the lock got an HTTP 500 where the
same request previously returned empty metrics. Reproduced before fixing:
"Expected response status code [200] but received 500". The controller now
degrades to empty numbers and still reports, which also covers the same 500
for any platform whose refresh 5xx'd — possible before this branch too.

The fallback verify was throwing away a result worth keeping. It is a billed
call on X and it proves the token alive exactly as a refresh does, so the
pre-publish check was paying to ask the same question minutes later. Stamped
like the other two sites.

Also rewrites a comment in rotatedTokenFrom() that described the data_get()
call it replaced rather than the blank() check beneath it. The review also
reported a false @throws on that method; it has no docblock at all.

Both fixes checked by reverting them and confirming the new tests fail.

* Degrade analytics on an unreachable platform, not on a bug

The rescue() added last commit caught Throwable, so it did not just absorb a
platform being down — it absorbed everything. A TypeError in any metrics
service rendered as "this account has no activity", with a log line as the only
sign anything was wrong. Reproduced: a metrics service throwing RuntimeException
returned 200 with empty metrics.

This is the same mistake the review flagged two rounds ago in
accessTokenStillWorks(), where treating any exception as "the token is healthy"
hid real failures. Narrowed the same way: PlatformUnavailableException and
ConnectionException degrade to empty numbers and still report, everything else
surfaces as the 500 it is.

The match moved into a named method so the intent has somewhere to live, since
the reason for the narrow catch matters more than the catch itself.

Both directions are pinned: widening the catch back to Throwable fails the
bug-is-not-hidden test, and removing the degradation fails the lock-collision
test.

* Trim the commentary back to what the code cannot say

RefreshSocialToken had 72 comment lines against 95 of code — 43% of the file.
The rest of the branch was heading the same way: two constants in
ConnectionVerifier carried twelve- and eight-line docblocks, and
VERIFIED_WITHIN_HOURS had twelve lines explaining a number.

Most of it was history rather than reasoning: what the code used to do, which
review round asked for a change, the full argument for a decision the code
already states. Kept the parts a reader cannot recover — why a catch is narrow,
why a stamp is not written inside refreshToken(), why the lock has to outlast
the timeouts — and cut the rest.

shouldTrustAWorkingAccessToken() went with it: a one-line method behind a
twelve-line docblock, now the condition it wrapped, inline where it is used.

No behaviour change; full suite unchanged at 3786.
2026-08-20 17:58:16 -03:00
Paulo Castellano
097f30c765
Fix 500 on the MCP settings page when a client has never been used (#297)
* Fix 500 on the MCP settings page when a client has never been used

Sorting the connected clients used the raw last_used_at value with an
int fallback, so a workspace holding both a used client (CarbonImmutable)
and a freshly connected one (null, falling back to 0) handed arsort() a
mix of object and int. PHP cannot compare the two and raises
"Object of class Carbon\CarbonImmutable could not be converted to int",
which Laravel's error handler turns into a 500.

Sorting on the timestamp keeps the key a single type; never-used clients
get 0 and land at the bottom, which is the intended order.

The existing ordering test only covered clients that had all been used,
which is why this shipped. The new test covers the mixed case.

* Cover the all-never-used case on the MCP clients list
2026-08-19 09:52:54 -03:00
Paulo Castellano
0eff22dc0d
Remove the post-templates feature (#296)
* Remove the post-templates feature

Removes the browsable post-templates catalog end to end: routes
(app.post-templates.index/apply), controller, form requests, resource,
Registry/PostTemplateData/TemplateNotFoundException services, the
templates:report console command, the templates/ catalog (36 files
across en/es/pt-BR), the templates Index.vue page, the template card
on the post Create.vue choice screen, and its i18n keys across all 16
locales.

Also removes TemplateContextResolver, which fed catalog examples into
PostContentGenerator's AI prompt. The "examples" block (with its
heading) is removed from generator.blade.php rather than left empty;
AI generation no longer references the templates catalog.

The AI content-template system under app/Ai/Templates (image/tweet
card templates for the AI post wizard) and TemplateImageGenerator are
unrelated and untouched.

* Update FormRequest naming examples after the post-templates removal

CLAUDE.md and the Cursor rules cited ApplyPostTemplateRequest and
IndexPostTemplateRequest as naming examples. Both classes were deleted
with the feature, so the guidance pointed at files that no longer exist.

* Share the AiTemplate type between the create screen and the AI wizard

Create.vue and AiPostWizard.vue each declared their own AiTemplate
interface, and they had drifted: the wizard's carried
applies_brand_visuals, the create screen's did not. TypeScript treated
them as unrelated types with the same name, which surfaced as a TS2719
error where the templates prop is passed between them.

The backend sends seven fields (PostController::create), so the shared
declaration follows the payload rather than the union of the two copies.
2026-08-19 09:34:26 -03:00